Biological databases are organized collections of biological information such as DNA, RNA, proteins, genes, genomes, and protein structures. They help researchers and students store, search, compare, and analyze biological data easily.
In this article, we will learn about the types of biological databases, important examples, and their uses in bioinformatics.
1. What Are Biological Databases?
Biological databases are online libraries where scientists store, share, and search biological data. This data can be DNA and RNA sequences, protein sequences, 3D protein structures, metabolic pathways, gene functions, or published research papers.
Every day, laboratories around the world produce a huge amount of biological data. Without well-organised databases, this data would be scattered, hard to find and almost impossible to reuse. Databases solve this problem by keeping the data in one place, in a standard format, so that anyone can search it, download it and build on it.
Daily-life example: Like Google Drive or your phone's gallery, where files are saved in one place and found with a search.
Biological databases are the backbone of bioinformatics, the field that uses computers to understand biological information. If you are a student of biotechnology, microbiology, genetics, biochemistry, pharmacy, or medicine, knowing these databases is a basic skill.
2. Why Do Biological Databases Matter?
Biological databases are useful because they:
- Save time and money. Researchers can reuse existing data instead of repeating experiments.
- Make science open. Most major databases are free for everyone to use.
- Help in discovery. Comparing sequences or structures can reveal the function of an unknown gene or protein.
- Support medicine. They are used to study diseases, find drug targets and track mutations.
- Keep data standard. Shared formats and identifiers (called accession numbers) make it easy to refer to the same record anywhere in the world.
3. Types of Biological Databases
Biological databases are grouped based on how the data is collected and processed. The three main groups are primary, secondary, and composite databases.
Primary Databases
Primary databases hold raw data submitted directly by researchers. Little or no processing is done, and the data is stored as submitted. The submitter is responsible for the data they upload. Examples: GenBank (DNA), ENA, DDBJ, and PDB (structures).
Daily-life example: The raw, unedited photos in your phone's gallery, stored exactly as you clicked them.
Secondary Databases
Secondary databases are built from primary data after analysis and curation. Experts or software add useful details such as families, patterns, and classifications. The information is more organised and reliable than raw data. Examples: Swiss-Prot, Pfam, and PROSITE.
Daily-life example: A photo album made from those photos, where the best ones are picked, edited, and given captions.
Composite Databases
Composite databases combine data from several primary and secondary sources in one place. They save time because you can search many sources at once. Example: UniProt.
Daily-life example: A price-comparison app that shows products from many online shops on one screen.
Primary vs Secondary vs Composite at a glance
| Feature | Primary | Secondary | Composite |
|---|---|---|---|
| Source of data | Directly from researchers | Derived from primary data | Merged from many databases |
| Processing | Minimal | Analysed and curated | Integrated |
| Reliability | Depends on the submitter | Higher | Depends on the sources included |
| Examples | GenBank, PDB | Swiss-Prot, Pfam, PROSITE | UniProt |
4. Nucleotide (DNA/RNA) Databases
Nucleotide databases store DNA and RNA sequences from different types of organisms, ranging from bacteria to humans. These databases help researchers find, compare, and analyse nucleotide sequences for studying genes, genomes, mutations, and other biological processes.
NCBI (National Center for Biotechnology Information)
NCBI is a large US-based hub that offers many databases and tools in one place. It is home to GenBank and PubMed, and also hosts resources such as RefSeq, Gene, dbSNP, ClinVar and the BLAST search tool.
Daily-life example: A big public library with many sections, where one building holds books, journals and computers.
GenBank, ENA (EMBL) and DDBJ
These three are the world's main public nucleotide sequence archives.
- GenBank: NCBI's public DNA sequence collection (USA).
- ENA (European Nucleotide Archive): The modern European sequence database, historically managed by and referred to as EMBL.
- DDBJ (DNA Data Bank of Japan): The Japanese nucleotide database.
The three share data regularly through the International Nucleotide Sequence Database Collaboration (INSDC), so they hold nearly the same records. A sequence submitted to one of them becomes available in all three.
Daily-life example: The same news story reported by three agencies in different countries. Each publishes it, so you get the same story from any of them.
Tip: Use GenBank when you want a sequence record along with its annotations, such as gene location, source organism and references.
5. Protein Sequence Databases
Protein sequence databases store the amino acid sequences of proteins along with details about their function, location in the cell and related diseases.
UniProt
UniProt is a central, composite resource for protein sequences and their functions. It is one of the most widely used protein databases in the world, and it brings together two main sections, Swiss-Prot and TrEMBL, under the UniProt Knowledgebase (UniProtKB).
Daily-life example: Wikipedia, a single place where you can look up almost anything and find its details.
Swiss-Prot
Entries are reviewed and checked by expert curators. The database provides high-quality and reliable information. It is smaller in size because every entry requires manual effort.
Daily-life example: A health article that a doctor has checked before publishing, so you can trust it.
TrEMBL
Entries are annotated automatically by computers. It is much larger in size, but the entries are not manually reviewed. Some entries may later be reviewed and moved to Swiss-Prot.
Daily-life example: The auto-tags your phone puts on photos, such as 'beach' or 'food'. They are quick and plentiful, but sometimes wrong.
| Feature | Swiss-Prot | TrEMBL |
|---|---|---|
| Annotation | Manual (expert-reviewed) | Automatic (computer-generated) |
| Quality | Very high | Good, but not verified |
| Size | Smaller | Much larger |
| Best for | Reliable, well-studied proteins | Wider coverage, newly sequenced proteins |
6. Protein Structure Databases
A protein's function depends on its 3D shape. Structure databases store these shapes so that scientists can study how proteins work and design drugs against them.
PDB (Protein Data Bank)
The PDB stores 3D structures of proteins and other biomolecules such as DNA, RNA and complexes. The structures are determined by experimental methods like X-ray crystallography, NMR spectroscopy and cryo-electron microscopy. The data is managed by the worldwide Protein Data Bank (wwPDB) partnership.
Daily-life example: The 3D view on a shopping app, where you rotate a sofa to see its exact shape before buying.
Good to know: Predicted structures, for example from AlphaFold, are available through the AlphaFold Protein Structure Database. These are computer predictions and are different from experimentally solved PDB structures.
SCOP / SCOP2 and CATH
These databases classify protein structures by shape (fold) and family. They help us understand how structure relates to function. SCOP and CATH use different approaches, so comparing both gives a fuller picture.
Note: The classic SCOP system has evolved into SCOP2 to better handle the massive, complex modern influx of structural data.
Daily-life example: Aisles in a supermarket, where items are sorted by type, shape and use so you can find them quickly.
7. Protein Family and Motif Databases
Proteins that share an ancestor or a job often look alike in their sequence. These databases group proteins by such similarities, which helps predict the function of new proteins.
Pfam
Pfam groups proteins into families and domains based on shared features, using statistical models called hidden Markov models (HMMs). Pfam data is now accessed through InterPro.
💡 Daily-life example: Music apps sorting songs into playlists by genre, so songs with similar features sit together.
PROSITE
PROSITE records short patterns and motifs that signal a protein's function, such as an enzyme's active site or a binding region.
Daily-life example: A barcode on a product. A short pattern tells the scanner exactly what the item is.
InterPro
InterPro brings together several family, domain and motif databases (including Pfam and PROSITE) for easy searching. You paste a protein sequence, and it shows all the matching families and domains in one result.
Daily-life example: A search engine that checks many review websites at once and shows you one combined result.
8. Pathway Databases
Genes and proteins rarely work alone. They work together in chains of reactions called pathways. Pathway databases show these connections.
KEGG (Kyoto Encyclopedia of Genes and Genomes)
KEGG maps pathways showing how genes, proteins and chemicals work together, such as metabolism, signalling and disease pathways.
💡 Daily-life example: A metro map, where lines and stations show how you get from one place to another.
Reactome
Reactome provides expert-reviewed, peer-reviewed details of biological pathways and reactions, and it is free and open to use.
Daily-life example: A step-by-step recipe in a cooking app, showing each stage from ingredients to the finished dish.
9. Genome Browsers
A genome is very long, with billions of letters in human DNA. Genome browsers let you view and explore genomes visually, zooming from a whole chromosome down to a single gene.
Ensembl
Ensembl lets you view and explore annotated genomes of many species. You can see genes, transcripts, variants and comparisons between species.
Daily-life example: Google Maps, where you can pick a place, zoom in and see every road and landmark labelled.
UCSC Genome Browser
The UCSC Genome Browser is a visual tool to browse genes and genome features along a chromosome. It uses 'tracks' that you can switch on and off.
Daily-life example: Switching Google Maps layers (traffic, satellite, terrain) to see different details of the same road.
10. Literature Database: PubMed
PubMed
PubMed is a free search engine for biomedical research papers. It is run by the US National Library of Medicine at NCBI and mainly covers the MEDLINE database and other life-science journals. It is useful for finding the science behind any gene, protein or disease.
Daily-life example: Searching online before a doctor's visit to read research on your symptoms or a medicine.
Search tip: Use keywords with AND, OR and NOT (for example, BRCA1 AND breast cancer) to narrow your results.
11. Other Useful Databases
Once you are comfortable with the basics, these databases will make your work easier:
- RefSeq (NCBI): A curated, non-redundant set of reference sequences for genomes, transcripts and proteins.
- dbSNP and ClinVar: Catalogues of genetic variants and their links to health conditions.
- GEO (Gene Expression Omnibus): A public archive of gene expression datasets.
- Gene Ontology (GO): A standard vocabulary to describe gene functions.
- STRING: Shows known and predicted protein–protein interactions.
- OMIM: A catalogue of human genes and genetic disorders.
12. Quick Comparison Table
| Category | Database | What it stores | Type |
|---|---|---|---|
| Nucleotide | GenBank, ENA, DDBJ | DNA/RNA sequences | Primary |
| Protein sequence | UniProt (Swiss-Prot, TrEMBL) | Protein sequences and functions | Composite |
| Protein structure | PDB | 3D structures | Primary |
| Structure classification | SCOP/SCOP2, CATH | Fold and family classification | Secondary |
| Family and motif | Pfam, PROSITE, InterPro | Families, domains, patterns | Secondary / Integrated |
| Pathway | KEGG, Reactome | Pathways and reactions | Secondary |
| Genome browser | Ensembl, UCSC | Annotated genomes | Browser/Resource |
| Literature | PubMed | Research papers | Literature |
13. How to Choose the Right Database
Ask yourself what you want to find, then pick the matching resource:
| If you want to... | Use |
|---|---|
| Find a DNA or RNA sequence | GenBank, ENA or DDBJ |
| Find a reliable protein sequence and its function | UniProt (Swiss-Prot) |
| See the 3D shape of a protein | PDB |
| Understand which family or domain a protein belongs to | Pfam, InterPro |
| Find a short functional pattern | PROSITE |
| See how genes work together in a process | KEGG or Reactome |
| Explore a gene on a chromosome | Ensembl or UCSC |
| Read research papers | PubMed |
14. A Simple Research Workflow Using These Databases
Here is how the databases connect in a real mini-project. Suppose you want to study the human gene TP53.
- Find the gene sequence. Search GenBank (or NCBI Gene) to get the DNA sequence and its accession number.
- Get the protein details. Search UniProt and open the reviewed Swiss-Prot entry for p53 to see its function and domains.
- Check its structure. Use the PDB to view the 3D structure of the protein.
- Identify domains and motifs. Paste the protein sequence into InterPro to see families, domains, and motifs.
- Study the pathway. Search KEGG or Reactome to see how the protein fits into pathways such as the cell cycle and apoptosis.
- View it in the genome. Open Ensembl or UCSC to see where the gene sits on its chromosome.
- Read the research. Search PubMed for recent papers on the gene and related diseases.
This is the general logic of bioinformatics: go from sequence, to structure, to function, to pathway, to literature.
15. Common Mistakes Beginners Make
- Treating all entries as equally reliable. Prefer reviewed entries (Swiss-Prot) when accuracy matters.
- Ignoring accession numbers. Always note the accession ID so others can find the exact same record.
- Confusing predicted and experimental structures. Check how a structure was obtained before drawing conclusions.
- Using only one database. Cross-checking across databases gives a more complete and trustworthy answer.
- Not checking the date. Databases are updated regularly, so note the version or date you used.
16. Frequently Asked Questions (FAQs)
What is a biological database?
A biological database is an organised online collection of biological data, such as sequences, structures, pathways or research papers, that scientists can search and share.
What are the main types of biological databases?
The three main types are primary databases (raw data, such as GenBank and PDB), secondary databases (curated data, such as Swiss-Prot, Pfam and PROSITE) and composite databases (combined data, such as UniProt).
What is the difference between primary and secondary databases?
Primary databases store raw data submitted directly by researchers. Secondary databases are built from primary data after analysis, adding details such as families, patterns, and classifications.
What is the difference between Swiss-Prot and TrEMBL?
Swiss-Prot entries are manually reviewed by experts and are highly reliable. TrEMBL entries are annotated automatically by computers, so the database is larger but not manually checked.
What are GenBank, ENA and DDBJ?
They are the three major public DNA sequence databases, located in the USA, Europe, and Japan. They exchange data through INSDC, so they hold nearly the same records.
Which database is used for protein 3D structures?
The Protein Data Bank (PDB) is the main database for experimentally determined 3D structures of proteins and other biomolecules.
What is the difference between KEGG and Reactome?
KEGG gives broad maps of pathways, including metabolism and disease, while Reactome offers detailed, expert-reviewed descriptions of biological reactions and pathways.
Are biological databases free?
Most major databases, including GenBank, UniProt, PDB, Ensembl, UCSC, Reactome and PubMed, are free for academic and general use. Some, like parts of KEGG, have licensing conditions for bulk or commercial use, so check each site's terms.
Which database should a beginner start with?
Start with NCBI (GenBank and PubMed), UniProt and PDB. These three cover sequences, protein information and structures, which is the foundation for most bioinformatics work.
17. Conclusion
Biological databases make it easier to find, analyse, and understand biological data. Each database has a specific purpose, from storing raw sequences to providing curated information about proteins, structures, genes, and pathways.
Understanding the purpose of databases such as GenBank, UniProt, PDB, Pfam, KEGG, Ensembl, and PubMed helps bioinformatics students choose the right resource for their work. With regular practice, searching and using these databases becomes easier and more useful for biological research.