A bioinformatics file format is a standard way of storing biological data so that it can be read and processed by different software and tools. Different formats are used for different types of data, such as DNA sequences, sequencing reads, gene annotations, genetic variants, and protein structures.
In this article, we will learn about the most commonly used bioinformatics file formats, their structure, common uses, and important differences between them.
FASTA File Format
FASTA is a simple text-based file format used to store DNA, RNA, or protein sequences. It is commonly used because its structure is easy to read and process.
A FASTA record has two main parts:
- Header: Starts with the > symbol and contains the sequence identifier and optional description.
- Sequence: Contains the DNA, RNA, or protein sequence.
For example:
>sample_gene1 Homo sapiens example gene ATGGCGTACGTTAGCTAGCTAGGCTAACGT TTAGCCGATCGATCGTAGCTAGCTAGCTAA
Here, sample_gene1 is the sequence identifier. The text after the identifier provides additional information about the sequence.
A FASTA file can contain one or many sequences. Long sequences can be written across multiple lines, and programs read them as a single continuous sequence.
Common extensions: .fa, .fasta, .fna for nucleotide sequences and .faa for protein sequences.
Common uses: FASTA files are used for reference genomes, gene sequences, protein sequences, and sequence searches such as BLAST.
Important: FASTA does not contain sequencing quality scores. If quality information is required, FASTQ is normally used.
FASTQ File Format
FASTQ is a text-based file format that stores sequencing reads along with a quality score for each base. It is commonly produced by DNA and RNA sequencing machines.
Each FASTQ record contains four lines:
@read1 GATTACAGGT + IIIIHHGG5#
- Line 1: Starts with @ and contains the read identifier.
- Line 2: Contains the sequence.
- Line 3: Starts with + and may optionally contain the read identifier.
- Line 4: Contains the quality scores for the sequence.
The sequence and quality string must have the same number of characters.
What is a Phred Quality Score?
A Phred quality score shows the probability that a sequencing base has been called incorrectly.
The formula is:
Q = −10 × log₁₀(P)
Here, P is the probability that the base call is wrong.
| Phred Score | Error Probability | Accuracy |
|---|---|---|
| Q10 | 1 in 10 | 90% |
| Q20 | 1 in 100 | 99% |
| Q30 | 1 in 1,000 | 99.9% |
| Q40 | 1 in 10,000 | 99.99% |
A higher Phred score means a lower probability of an incorrect base call.
GenBank File Format
GenBank is a sequence record format used by NCBI to store biological sequences together with detailed information and annotations.
Unlike FASTA, which mainly stores sequence information, a GenBank record can contain information about the organism, sequence, genes, coding regions, references, and other features.
A simplified GenBank record looks like this:
LOCUS DEMO0001 60 bp DNA linear
DEFINITION Example bacterial gene for teaching.
ACCESSION DEMO0001
VERSION DEMO0001.1
SOURCE Escherichia coli
ORGANISM Escherichia coliFEATURES Location/Qualifiers
source 1..60
/organism="Escherichia coli"
gene 1..51
/gene="demA"
CDS 1..51
/gene="demA"
/product="demonstration protein"ORIGIN
1 atggcgtacg ttagctagct aggctaacgt
31 ttagccgatc gatcgtagct agctagctaa
//
Important parts include:
- LOCUS: Provides basic information about the record.
- ACCESSION: Provides the record's accession identifier.
- FEATURES: Contains annotations such as genes and coding sequences.
- ORIGIN: Contains the biological sequence.
- // marks the end of the record.
A location such as 1..51 means that the feature starts at position 1 and ends at position 51.
Common use: GenBank is useful when you need both the biological sequence and its associated annotations.
GFF and GTF File Formats
GFF (General Feature Format) and GTF (Gene Transfer Format) are text-based formats used to store genomic annotations such as genes, transcripts, and exons.
They usually store annotation separately from the actual genome sequence.
A GFF3 file uses nine columns:
- Sequence name
- Source
- Feature type
- Start
- End
- Score
- Strand
- Phase
- Attributes
For example:
chr1 demo gene 1000 2200 . + . ID=gene1;Name=demA
chr1 demo mRNA 1000 2200 . + . ID=tx1;Parent=gene1
chr1 demo exon 1000 1300 . + . ID=ex1;Parent=tx1
The Parent attribute shows the relationship between features. For example, an exon can belong to a transcript, and the transcript can belong to a gene.
GTF uses a similar structure but has different rules for its attribute field. A common GTF record looks like this:
chr1 demo exon 1000 1300 . + . gene_id "gene1"; transcript_id "tx1";
Important: GFF and GTF coordinates are 1-based and inclusive. This means both the start and end positions are included.
Common uses: GFF and GTF files are widely used for genome annotation, gene prediction, RNA-seq analysis, and transcript analysis.
BED File Format
BED (Browser Extensible Data) is a simple text format used to describe regions of a genome.
The first three BED columns are:
- Chromosome
- Start position
- End position
For example:
chr1 100 200 chr1 450 520 chr2 900 1000
BED can also contain additional columns such as a name, score, and strand.
Common uses: BED files are used for genomic regions such as genes, exons, sequencing peaks, and other regions of interest. They are also commonly used with genome browsers and tools such as bedtools.
Important: BED uses 0-based, half-open coordinates. The start position is included, but the end position is not.
Difference Between 0-Based and 1-Based Coordinates
Coordinate systems are important when working with genomic data because different formats use different systems.
- GFF/GTF: 1-based and inclusive
- BED: 0-based and half-open
- SAM: Alignment position (POS) is 1-based
- VCF: Variant position (POS) is 1-based
For example, consider the sequence:
ACGTACGTAC
If we want the bases from position 3 to position 6:
1 2 3 4 5 6 7 8 9 10 A C G T A C G T A C
In GFF/GTF, the region is written as:
3 6
In BED, the same region is written as:
2 6
The BED length is calculated as:
end − start = 6 − 2 = 4
This coordinate difference is important when converting or comparing genomic files.
SAM and BAM File Formats
SAM (Sequence Alignment/Map) is a text format used to store sequencing reads aligned to a reference sequence.
BAM stores the same type of alignment information in a compressed binary format. BAM files are smaller and faster to process than SAM files.
A SAM file contains header lines followed by alignment records.
A simplified alignment record looks like this:
read1 99 chr1 1000 60 10M = 1050 60 GATTACAGGT IIIIHHGG5# NM:i:0
The main SAM columns include:
| Column | Meaning |
|---|---|
| QNAME | Read name |
| FLAG | Information about the read and its alignment |
| RNAME | Reference sequence name |
| POS | Alignment start position |
| MAPQ | Mapping quality |
| CIGAR | Describes the alignment |
| RNEXT | Reference name of the mate |
| PNEXT | Position of the mate |
| TLEN | Template length |
| SEQ | Read sequence |
| QUAL | Base quality scores |
CIGAR String
The CIGAR string describes how a read is aligned to the reference.
For example:
10M
means that 10 read bases are aligned using the M operation.
Other common CIGAR operations include:
- I - Insertion
- D - Deletion
- S - Soft clipping
- N - Skipped region
For example:
5M2I3M
means 5 aligned bases, followed by a 2-base insertion, followed by 3 aligned bases.
SAM FLAG
The FLAG field stores information about the alignment using a set of numeric values.
Some common values are:
| Value | Meaning |
|---|---|
| 1 | Read is paired |
| 2 | Read is part of a proper pair |
| 4 | Read is unmapped |
| 16 | Read is mapped to the reverse strand |
| 32 | Mate is mapped to the reverse strand |
| 64 | First read in a pair |
| 128 | Second read in a pair |
| 1024 | Read is a duplicate |
For example, FLAG 99 is made from:
1 + 2 + 32 + 64 = 99
Tools such as samtools can be used to read and interpret SAM/BAM files.
Basic samtools Commands
samtools view -b sample.sam > sample.bam samtools sort sample.bam -o sorted.bam samtools index sorted.bam samtools flagstat sorted.bam
These commands can be used to convert SAM to BAM, sort a BAM file, create an index, and generate a basic alignment summary.
VCF File Format
VCF (Variant Call Format) is a text-based format used to store genetic variants identified in a sample.
A variant can be a single-base change, insertion, deletion, or another type of genetic difference.
A VCF file contains header lines followed by variant records.
For example:
##fileformat=VCFv4.2
##INFO=<ID=DP,Number=1,Type=Integer,Description="Total depth">
#CHROM POS ID REF ALT QUAL FILTER INFO FORMAT SAMPLE1
chr1 12345 rs123 A G 50 PASS DP=30 GT:DP 0/1:28
chr1 20400 . AT A 42 PASS DP=25 GT:DP 1/1:25
Important VCF columns include:
- CHROM: Chromosome or reference sequence
- POS: Position of the variant
- REF: Reference allele
- ALT: Alternate allele
- QUAL: Variant quality score
- FILTER: Filtering status
- INFO: Additional information
- FORMAT: Describes sample-specific fields
Genotype Values
The GT field describes the genotype of a sample.
| Genotype | Meaning |
|---|---|
| 0/0 | Both alleles match the reference |
| 0/1 | One reference allele and one alternate allele |
| 1/1 | Both alleles are alternate |
| 0|1 | Phased genotype |
For example:
chr1 12345 rs123 A G
means that the reference allele is A and the alternate allele is G.
VCF files are commonly compressed and indexed for faster processing. Tools such as bcftools, bgzip, and tabix are widely used with VCF files.
PDB and mmCIF File Formats
PDB and PDBx/mmCIF are file formats used to store information about the three-dimensional structures of proteins, nucleic acids, and other biological molecules.
Structure files contain information such as atom names, residue names, chains, and the three-dimensional coordinates of atoms.
PDB Format
The PDB format is a fixed-width text format. Important records include ATOM and HETATM.
For example:
ATOM 1 N MET A 1 27.340 24.430 2.614 1.00 11.99 N
ATOM 2 CA MET A 1 26.266 25.413 2.842 1.00 11.85 C
HETATM 900 O HOH A 301 15.100 10.220 7.400 1.00 20.10 O
Important information includes:
- Atom number
- Atom name
- Residue name
- Chain identifier
- Residue number
- X, Y, and Z coordinates
- Occupancy
- B-factor
- Element
The X, Y, and Z coordinates are usually measured in ångströms (Å).
PDBx/mmCIF
PDBx/mmCIF is the modern format used for representing structural data in the Protein Data Bank. Unlike the fixed-width PDB format, mmCIF stores information using named fields and tables.
A simple example is:
loop_
_atom_site.group_PDB
_atom_site.id
_atom_site.label_atom_id
_atom_site.label_comp_id
_atom_site.label_asym_id
_atom_site.Cartn_x
_atom_site.Cartn_y
_atom_site.Cartn_z
ATOM 1 N MET A 27.340 24.430 2.614
ATOM 2 CA MET A 26.266 25.413 2.842
The loop_ keyword starts a table, followed by the names of the fields and their values.
mmCIF is better suited for large and complex structures because it does not have the same field and identifier limitations as the older PDB format.
Common Conversion Tools
Bioinformatics workflows often require converting data from one format to another. Several command-line tools and libraries are available for this purpose.
| Task | Tool | Example |
|---|---|---|
| FASTQ to FASTA | seqkit | seqkit fq2fa reads.fq -o reads.fa |
| FASTQ to FASTA | seqtk | seqtk seq -a reads.fq > reads.fa |
| GenBank to FASTA | Biopython | SeqIO.convert() |
| GFF3 to GTF | gffread | gffread in.gff3 -T -o out.gtf |
| GTF to GFF3 | gffread | gffread in.gtf -o out.gff3 |
| BAM to BED | bedtools | bedtools bamtobed -i sorted.bam |
| BED regions to FASTA | bedtools | bedtools getfasta -fi genome.fa -bed regions.bed |
| SAM to BAM | samtools | samtools view -b in.sam > out.bam |
| BAM to SAM | samtools | samtools view -h in.bam > out.sam |
| BAM to FASTQ | samtools | samtools fastq sorted.bam > reads.fq |
| VCF to BCF | bcftools | bcftools view -O b in.vcf -o out.bcf |
| Compress VCF | bgzip | bgzip in.vcf |
| Index VCF | tabix | tabix -p vcf in.vcf.gz |
Typical Bioinformatics Workflow
Different file formats are often used at different stages of a bioinformatics workflow.
A basic variant analysis workflow can look like this:
- FASTQ files contain sequencing reads from the sequencing machine.
- Quality control is performed on the reads.
- The reads are aligned to a reference FASTA file.
- The alignment is stored in SAM/BAM format.
- A variant caller identifies variants and produces a VCF file.
- GFF/GTF or BED files can be used to identify genes and genomic regions related to the variants.
- If a variant affects a protein-coding gene, its protein structure can be studied using PDB/mmCIF data.
Quick Comparison of Bioinformatics File Formats
| Format | Main purpose | Coordinate system |
|---|---|---|
| FASTA | DNA, RNA, or protein sequences | Not applicable |
| FASTQ | Sequences with quality scores | Not applicable |
| GenBank | Sequences with detailed annotations | 1-based |
| GFF/GTF | Gene and feature annotations | 1-based, inclusive |
| BED | Genomic regions | 0-based, half-open |
| SAM/BAM | Aligned sequencing reads | SAM POS is 1-based |
| VCF | Genetic variants | 1-based |
| PDB/mmCIF | 3D molecular structures | X, Y, Z coordinates |
Conclusion
Bioinformatics file formats are used to store different types of biological data. FASTA stores sequences, FASTQ stores sequences with quality scores, GenBank stores sequences with annotations, GFF/GTF and BED store genomic features and regions, SAM/BAM store alignments, VCF stores variants, and PDB/mmCIF store molecular structures.
Understanding the basic structure and purpose of these formats makes it easier to work with bioinformatics tools and analyze biological data.