Home › Bioinformatics Tutorial › Biological Databases

Biological Databases

⏱ 13 min read Updated: 08 Oct 2026

Biological databases are organized collections of biological information such as DNA, RNA, proteins, genes, genomes, and protein structures. They help researchers and students store, search, compare, and analyze biological data easily.

In this article, we will learn about the types of biological databases, important examples, and their uses in bioinformatics.

1. What Are Biological Databases?

Biological databases are online libraries where scientists store, share, and search biological data. This data can be DNA and RNA sequences, protein sequences, 3D protein structures, metabolic pathways, gene functions, or published research papers.

Every day, laboratories around the world produce a huge amount of biological data. Without well-organised databases, this data would be scattered, hard to find and almost impossible to reuse. Databases solve this problem by keeping the data in one place, in a standard format, so that anyone can search it, download it and build on it.

Daily-life example: Like Google Drive or your phone's gallery, where files are saved in one place and found with a search.

Biological databases are the backbone of bioinformatics, the field that uses computers to understand biological information. If you are a student of biotechnology, microbiology, genetics, biochemistry, pharmacy, or medicine, knowing these databases is a basic skill.

2. Why Do Biological Databases Matter?

Biological databases are useful because they:

  • Save time and money. Researchers can reuse existing data instead of repeating experiments.
  • Make science open. Most major databases are free for everyone to use.
  • Help in discovery. Comparing sequences or structures can reveal the function of an unknown gene or protein.
  • Support medicine. They are used to study diseases, find drug targets and track mutations.
  • Keep data standard. Shared formats and identifiers (called accession numbers) make it easy to refer to the same record anywhere in the world.

3. Types of Biological Databases

Biological databases are grouped based on how the data is collected and processed. The three main groups are primary, secondary, and composite databases.

Primary Databases

Primary databases hold raw data submitted directly by researchers. Little or no processing is done, and the data is stored as submitted. The submitter is responsible for the data they upload. Examples: GenBank (DNA), ENA, DDBJ, and PDB (structures).

Daily-life example: The raw, unedited photos in your phone's gallery, stored exactly as you clicked them.

Secondary Databases

Secondary databases are built from primary data after analysis and curation. Experts or software add useful details such as families, patterns, and classifications. The information is more organised and reliable than raw data. Examples: Swiss-Prot, Pfam, and PROSITE.

Daily-life example: A photo album made from those photos, where the best ones are picked, edited, and given captions.

Composite Databases

Composite databases combine data from several primary and secondary sources in one place. They save time because you can search many sources at once. Example: UniProt.

Daily-life example: A price-comparison app that shows products from many online shops on one screen.

Primary vs Secondary vs Composite at a glance

FeaturePrimarySecondaryComposite
Source of dataDirectly from researchersDerived from primary dataMerged from many databases
ProcessingMinimalAnalysed and curatedIntegrated
ReliabilityDepends on the submitterHigherDepends on the sources included
ExamplesGenBank, PDBSwiss-Prot, Pfam, PROSITEUniProt

4. Nucleotide (DNA/RNA) Databases

Nucleotide databases store DNA and RNA sequences from different types of organisms, ranging from bacteria to humans. These databases help researchers find, compare, and analyse nucleotide sequences for studying genes, genomes, mutations, and other biological processes.

NCBI (National Center for Biotechnology Information)

NCBI is a large US-based hub that offers many databases and tools in one place. It is home to GenBank and PubMed, and also hosts resources such as RefSeq, Gene, dbSNP, ClinVar and the BLAST search tool.

Daily-life example: A big public library with many sections, where one building holds books, journals and computers.

GenBank, ENA (EMBL) and DDBJ

These three are the world's main public nucleotide sequence archives.

  • GenBank: NCBI's public DNA sequence collection (USA).
  • ENA (European Nucleotide Archive): The modern European sequence database, historically managed by and referred to as EMBL.
  • DDBJ (DNA Data Bank of Japan): The Japanese nucleotide database.

The three share data regularly through the International Nucleotide Sequence Database Collaboration (INSDC), so they hold nearly the same records. A sequence submitted to one of them becomes available in all three.

Daily-life example: The same news story reported by three agencies in different countries. Each publishes it, so you get the same story from any of them.

Tip: Use GenBank when you want a sequence record along with its annotations, such as gene location, source organism and references.

5. Protein Sequence Databases

Protein sequence databases store the amino acid sequences of proteins along with details about their function, location in the cell and related diseases.

UniProt

UniProt is a central, composite resource for protein sequences and their functions. It is one of the most widely used protein databases in the world, and it brings together two main sections, Swiss-Prot and TrEMBL, under the UniProt Knowledgebase (UniProtKB).

Daily-life example: Wikipedia, a single place where you can look up almost anything and find its details.

Swiss-Prot

Entries are reviewed and checked by expert curators. The database provides high-quality and reliable information. It is smaller in size because every entry requires manual effort.

Daily-life example: A health article that a doctor has checked before publishing, so you can trust it.

TrEMBL

Entries are annotated automatically by computers. It is much larger in size, but the entries are not manually reviewed. Some entries may later be reviewed and moved to Swiss-Prot.

   Daily-life example: The auto-tags your phone puts on photos, such as 'beach' or 'food'. They are quick and plentiful, but sometimes wrong.

FeatureSwiss-ProtTrEMBL
AnnotationManual (expert-reviewed)Automatic (computer-generated)
QualityVery highGood, but not verified
SizeSmallerMuch larger
Best forReliable, well-studied proteinsWider coverage, newly sequenced proteins

6. Protein Structure Databases

A protein's function depends on its 3D shape. Structure databases store these shapes so that scientists can study how proteins work and design drugs against them.

PDB (Protein Data Bank)

The PDB stores 3D structures of proteins and other biomolecules such as DNA, RNA and complexes. The structures are determined by experimental methods like X-ray crystallography, NMR spectroscopy and cryo-electron microscopy. The data is managed by the worldwide Protein Data Bank (wwPDB) partnership.

   Daily-life example: The 3D view on a shopping app, where you rotate a sofa to see its exact shape before buying.

Good to know: Predicted structures, for example from AlphaFold, are available through the AlphaFold Protein Structure Database. These are computer predictions and are different from experimentally solved PDB structures.

SCOP / SCOP2 and CATH

These databases classify protein structures by shape (fold) and family. They help us understand how structure relates to function. SCOP and CATH use different approaches, so comparing both gives a fuller picture.

Note: The classic SCOP system has evolved into SCOP2 to better handle the massive, complex modern influx of structural data.

  Daily-life example: Aisles in a supermarket, where items are sorted by type, shape and use so you can find them quickly.

7. Protein Family and Motif Databases

Proteins that share an ancestor or a job often look alike in their sequence. These databases group proteins by such similarities, which helps predict the function of new proteins.

Pfam

Pfam groups proteins into families and domains based on shared features, using statistical models called hidden Markov models (HMMs). Pfam data is now accessed through InterPro.

💡 Daily-life example: Music apps sorting songs into playlists by genre, so songs with similar features sit together.

PROSITE

PROSITE records short patterns and motifs that signal a protein's function, such as an enzyme's active site or a binding region.

   Daily-life example: A barcode on a product. A short pattern tells the scanner exactly what the item is.

InterPro

InterPro brings together several family, domain and motif databases (including Pfam and PROSITE) for easy searching. You paste a protein sequence, and it shows all the matching families and domains in one result.

   Daily-life example: A search engine that checks many review websites at once and shows you one combined result.

8. Pathway Databases

Genes and proteins rarely work alone. They work together in chains of reactions called pathways. Pathway databases show these connections.

KEGG (Kyoto Encyclopedia of Genes and Genomes)

KEGG maps pathways showing how genes, proteins and chemicals work together, such as metabolism, signalling and disease pathways.

💡 Daily-life example: A metro map, where lines and stations show how you get from one place to another.

Reactome

Reactome provides expert-reviewed, peer-reviewed details of biological pathways and reactions, and it is free and open to use.

   Daily-life example: A step-by-step recipe in a cooking app, showing each stage from ingredients to the finished dish.

9. Genome Browsers

A genome is very long, with billions of letters in human DNA. Genome browsers let you view and explore genomes visually, zooming from a whole chromosome down to a single gene.

Ensembl

Ensembl lets you view and explore annotated genomes of many species. You can see genes, transcripts, variants and comparisons between species.

   Daily-life example: Google Maps, where you can pick a place, zoom in and see every road and landmark labelled.

UCSC Genome Browser

The UCSC Genome Browser is a visual tool to browse genes and genome features along a chromosome. It uses 'tracks' that you can switch on and off.

   Daily-life example: Switching Google Maps layers (traffic, satellite, terrain) to see different details of the same road.

10. Literature Database: PubMed

PubMed

PubMed is a free search engine for biomedical research papers. It is run by the US National Library of Medicine at NCBI and mainly covers the MEDLINE database and other life-science journals. It is useful for finding the science behind any gene, protein or disease.

   Daily-life example: Searching online before a doctor's visit to read research on your symptoms or a medicine.

Search tip: Use keywords with AND, OR and NOT (for example, BRCA1 AND breast cancer) to narrow your results.

11. Other Useful Databases

Once you are comfortable with the basics, these databases will make your work easier:

  • RefSeq (NCBI): A curated, non-redundant set of reference sequences for genomes, transcripts and proteins.
  • dbSNP and ClinVar: Catalogues of genetic variants and their links to health conditions.
  • GEO (Gene Expression Omnibus): A public archive of gene expression datasets.
  • Gene Ontology (GO): A standard vocabulary to describe gene functions.
  • STRING: Shows known and predicted protein–protein interactions.
  • OMIM: A catalogue of human genes and genetic disorders.

12. Quick Comparison Table

CategoryDatabaseWhat it storesType
NucleotideGenBank, ENA, DDBJDNA/RNA sequencesPrimary
Protein sequenceUniProt (Swiss-Prot, TrEMBL)Protein sequences and functionsComposite
Protein structurePDB3D structuresPrimary
Structure classificationSCOP/SCOP2, CATHFold and family classificationSecondary
Family and motifPfam, PROSITE, InterProFamilies, domains, patternsSecondary / Integrated
PathwayKEGG, ReactomePathways and reactionsSecondary
Genome browserEnsembl, UCSCAnnotated genomesBrowser/Resource
LiteraturePubMedResearch papersLiterature

13. How to Choose the Right Database

Ask yourself what you want to find, then pick the matching resource:

If you want to...Use
Find a DNA or RNA sequenceGenBank, ENA or DDBJ
Find a reliable protein sequence and its functionUniProt (Swiss-Prot)
See the 3D shape of a proteinPDB
Understand which family or domain a protein belongs toPfam, InterPro
Find a short functional patternPROSITE
See how genes work together in a processKEGG or Reactome
Explore a gene on a chromosomeEnsembl or UCSC
Read research papersPubMed

14. A Simple Research Workflow Using These Databases

Here is how the databases connect in a real mini-project. Suppose you want to study the human gene TP53.

  1. Find the gene sequence. Search GenBank (or NCBI Gene) to get the DNA sequence and its accession number.
  2. Get the protein details. Search UniProt and open the reviewed Swiss-Prot entry for p53 to see its function and domains.
  3. Check its structure. Use the PDB to view the 3D structure of the protein.
  4. Identify domains and motifs. Paste the protein sequence into InterPro to see families, domains, and motifs.
  5. Study the pathway. Search KEGG or Reactome to see how the protein fits into pathways such as the cell cycle and apoptosis.
  6. View it in the genome. Open Ensembl or UCSC to see where the gene sits on its chromosome.
  7. Read the research. Search PubMed for recent papers on the gene and related diseases.

This is the general logic of bioinformatics: go from sequence, to structure, to function, to pathway, to literature.

15. Common Mistakes Beginners Make

  • Treating all entries as equally reliable. Prefer reviewed entries (Swiss-Prot) when accuracy matters.
  • Ignoring accession numbers. Always note the accession ID so others can find the exact same record.
  • Confusing predicted and experimental structures. Check how a structure was obtained before drawing conclusions.
  • Using only one database. Cross-checking across databases gives a more complete and trustworthy answer.
  • Not checking the date. Databases are updated regularly, so note the version or date you used.

16. Frequently Asked Questions (FAQs)

What is a biological database?

A biological database is an organised online collection of biological data, such as sequences, structures, pathways or research papers, that scientists can search and share.

What are the main types of biological databases?

The three main types are primary databases (raw data, such as GenBank and PDB), secondary databases (curated data, such as Swiss-Prot, Pfam and PROSITE) and composite databases (combined data, such as UniProt).

What is the difference between primary and secondary databases?

Primary databases store raw data submitted directly by researchers. Secondary databases are built from primary data after analysis, adding details such as families, patterns, and classifications.

What is the difference between Swiss-Prot and TrEMBL?

Swiss-Prot entries are manually reviewed by experts and are highly reliable. TrEMBL entries are annotated automatically by computers, so the database is larger but not manually checked.

What are GenBank, ENA and DDBJ?

They are the three major public DNA sequence databases, located in the USA, Europe, and Japan. They exchange data through INSDC, so they hold nearly the same records.

Which database is used for protein 3D structures?

The Protein Data Bank (PDB) is the main database for experimentally determined 3D structures of proteins and other biomolecules.

What is the difference between KEGG and Reactome?

KEGG gives broad maps of pathways, including metabolism and disease, while Reactome offers detailed, expert-reviewed descriptions of biological reactions and pathways.

Are biological databases free?

Most major databases, including GenBank, UniProt, PDB, Ensembl, UCSC, Reactome and PubMed, are free for academic and general use. Some, like parts of KEGG, have licensing conditions for bulk or commercial use, so check each site's terms.

Which database should a beginner start with?

Start with NCBI (GenBank and PubMed), UniProt and PDB. These three cover sequences, protein information and structures, which is the foundation for most bioinformatics work.

17. Conclusion

Biological databases make it easier to find, analyse, and understand biological data. Each database has a specific purpose, from storing raw sequences to providing curated information about proteins, structures, genes, and pathways.

Understanding the purpose of databases such as GenBank, UniProt, PDB, Pfam, KEGG, Ensembl, and PubMed helps bioinformatics students choose the right resource for their work. With regular practice, searching and using these databases becomes easier and more useful for biological research.