Practical - Exploring sequence database
The bioinformatics page contains a lot of data and can be quite long; don’t hesitate to use Ctrl + F to quickly find the information you need.
NCBI - Interrogation of databases by Entrez
The NCBI (National Center for Biotechnology Information) is a key resource for genomic data and bioinformatics tools. It hosts databases like GenBank, Nucleotide, PubMed, Taxonomy… In this first part of the practical, we will explore several of the databases hosted by the NCBI.
Go to the NCBI website.
Use access number
Select Nucleotide database on the NCBI website and search for the sequence NC_045512.
Exercise 1 What is this sequence? How long is it?
Exercise 2 Consult the Comment field. Which bank does this sequence come from? Is it reliable? What sequencing technology was used? What can you say about this entry compared with the initial GenBank entry (1st sequence of the SARS-Cov-2 genome)?
Look for the coding region of the N protein, also known as nucleocapsid phosphoprotein (you can find it by searching the scientific name on the page).
Exercise 3 What do you think of the composition of this protein? What problems might such a sequence pose in a similarity search?
Exercise 4 Display the genome sequence in fasta format. How does this format compare with the RefSeq format?
Search by annotation (Filters/Advanced)
Breast cancer is a common cancer influenced by genetic factors, including BRCA1 or BRCA2.. Another critical gene is MUC1, often overexpressed in breast cancers, where it promotes tumor growth and metastasis. This exercise will explore the nucleotide sequences of MUC1 in humans using the NCBI database.
Search in Nucleotide for sequences containing the MUC1 gene for Homo sapiens, using filters or the advanced search.
Exercise 5 Why are there several sequences in RefSeq? in Genbank?
Among all the results, one of the sequences is a fragment of the gene in a breast cancer patient. To find it more easily, search for KT152889 with the search bar.
Exercise 6 Quickly position the annotated gene elements on a schema. What is the position of the variation on the sequence? What does it correspond to?
Among the features, the source field provides a cross-reference to NCBI’s Taxonomy bank. Follow this link.
Taxonomy bank
Exercise 7 What human subspecies are listed in the Taxonomy bank?
Exercise 8 What other organisms are listed in the Homininae ?
Exercise 9 What orders are closest to the primate order? Which has the most available proteins?
Check the Protein box under the search bar to display the number of proteins for each group. By default, 3 sublevels are displayed, but you can change it to 1 to make the information easier to read. Don’t forget to click Display to update the page.
Exercise 10 How many genomes are available in the Genome Bank for Animals (Metazoa), Fungi, Viridiplantae and Eukaryotes? Comment.
Check the Datasets Genome box under the search bar to display the number of genomes for each group.
Protein sequence search in UniprotKB
In the second part of the practical, we will explore protein-related databases and resources, such as UniProt, InterPro, and Pfam.
UniProt is a comprehensive protein sequence and functional information database, which includes Swiss-Prot, a curated, manually annotated section, and TrEMBL, a collection of computationally annotated protein sequences.
Go to the UniprotKB bank.
Advanced search fields
Exercise 11 How many human protein sequences are listed in UniprotKB? in SwissProt?
Example of a SwissProt entry
Exercise 12 Search for the protein associated with the MUC1 gene in human. Look the results. Which is the best entry?
Open the entry you consider the most reliable.
Exercise 13 Has the existence of this protein been established (see Protein existence at the top of the page)?
Exercise 14 Compare the default format, the text format and the fasta format (available in Download).
Exercise 15 How many isoforms are described for this protein (see Sequences & Isoforms)? Align the isoforms and see how they differ. For that, click on align isoforms and launch the alignment.
Function - GO annotations
Exercise 16 The Go ontology can be used to describe proteins in 3 ways. Which are they?
Exercise 17 Follow the link to the entry “negative regulation of intrinsic apoptotic signaling pathway in response to DNA damage by p53 class mediator”. What is it? What type of GO is it? How often is it used to describe proteins?
Exercise 18 Are there any synonyms for the term GO?
Structure
Exercise 19 What can you say about the structures obtained experimentally for the protein?
Exercise 20 What can you say about the reliability of the AlphaFold prediction for this protein?
Tab External links (top of page)
Sequence databases
Exercise 21 Why are there so many more links to EMBL/GenBank/DDBJ than to RefSeq?
Families and domains in INTERPRO database
Display the graphical representation (View protein in InterPro).
Exercise 22 How many distinct protein domains are there in this protein?
Exercise 23 From which banks do the 3 signatures associated with this InterPro domain originate?
You can get more information by passing the mouse over the colored bars, to open the domain in the associated bank, click on the corresponding link on the right.
Exercise 24 Follow the link to the associated PFAM signature. Display the signature (click on the Profile HMM tab). Comment the signature.
In a new tab, search for the BRCA1_HUMAN protein in InterProScan.
Exercise 25 How many distinct protein domains are there in this protein?