Skip to content

Isoform nomenclature question #217

Description

@MonicaPoelchau-USDA

Hi EGAPx team! I have some questions about the isoform nomenclature rules that EGAPx uses in product attributes. I've noticed three types of isoform product suffixes (where 'N' is a number):

  1. isoform N (for mRNA/CDS features)
  2. -TN (for mRNA/CDS features)
  3. transcript variant XN (for transcript features)

Some observations about the suffixes:

  • In some cases, the number used in the 'isoform' suffix and the '-T' suffix is different for the same isoform.
  • For protein-coding isoforms, the transcript ID uses '-RN' instead of '-TN'.
  • For transcript features, different rules for isoform naming are used.

My questions are:

  1. Why are two separate suffixes used in the same product value for protein-coding isoforms?
  2. Why do the numbers used in the separate suffixes sometimes differ for the same isoform?
  3. Why use 'TN' for the protein-coding product suffix, 'RN' for the protein-coding transcript ID suffix, and 'XN' for transcript feature product suffixes?

Here are some examples from EGAPx runs (v0.4.1 and v0.5):

  • product=heat shock protein 60A-like isoform 1-T1
  • product=uncharacterized protein isoform 1-T2
  • product=synaptotagmin 1 isoform 1-T3
  • product=uncharacterized protein isoform 2-T3
  • product=cramped chromatin regulator%2C transcript variant X5

Thanks for your help with this!

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions