AI in Genomics Training Wiki

Step-by-step GetGenome workflows for sequence alignment, genome annotation, protein modelling and structure analysis.

Contents

This wiki collects the GetGenome step-by-step workflows used in the AI in Genomics training: aligning a sequence, annotating a genome, modelling a protein, visualising the model and comparing structures. Each chapter can be read on its own, and the glossary at the end explains the terms used throughout.

  1. 1

    Compare a DNA or protein sequence with millions of known sequences using NCBI BLAST, and read the results.

    4 figures · 5 min read

  2. 2

    Predict genes and assign functions to an assembly with Prokka on the command line (2.1), Prokka on Galaxy EU (2.2) or RAST (2.3).

    19 min read

  3. 3

    Model your protein of interest with the AlphaFold 3 Server and judge how far the model can be trusted.

    4 figures · 9 min read

  4. 4

    Install UCSF ChimeraX, open an AlphaFold 3 model and analyse its confidence and its interface.

    8 figures · 6 min read

  5. 5

    Search for similar structures with Foldseek and align several structures with FoldMason.

    2 figures · 5 min read

  6. 6

    Key terms used across the workflows, filterable by topic and listed from A to Z.

    119 terms in 11 topics

1Sequence alignment

NCBI BLAST · 4 figures · 5 min read

NCBI BLAST compares a DNA or protein sequence with millions of known sequences in the NCBI database and finds the most similar ones. This workflow guides you through choosing the right type of BLAST, running a search or an alignment, and reading the results.

1.1Choosing BLASTStep 1

Go to NCBI BLAST. The type of BLAST depends on what your sequence is and what you want to compare it with:

ProgramYour sequenceCompared withUse it to
blastnDNADNAFind similar genes, or check a PCR product or clone
blastpProteinProteinFind sequence-similar proteins and predict function
blastxDNA, translatedProteinFind which protein a DNA sequence may encode
tblastnProteinDNA, translatedFind genes encoding your protein in genomes
tblastxDNA, translatedDNA, translatedCompare distantly related DNA sequences

Translated programs read DNA in all six reading frames, so they find proteins even when you do not know the gene structure.

1.2Running BLASTStep 2

BLAST can be run in two ways:

Search a databaseAlign two or more sequences
QuestionWhich known sequences are similar to mine?How similar are these specific sequences?
Compared withA whole database, such as all proteins in NCBIOnly the sequences you paste or upload
HowDefault settingTick Align two or more sequences below the query box
ExampleIdentify an unknown protein from your genomeCompare your protein with a known version from UniProt

To align many sequences together (for example, for a phylogenetic tree), use a multiple alignment tool such as Clustal Omega or MAFFT instead: BLAST compares your query with each sequence separately.

  1. Enter your query. Paste your sequence (FASTA format or bare sequence) or an NCBI accession number, or upload a file.

    Figure 1.1. NCBI BLAST (blastp) search page for a database search: enter the query, leave “Align two or more sequences” unticked, choose the database and click BLAST.
    Figure 1.2. The same page set up to align your own sequences: tick “Align two or more sequences”, then enter the query and the subject sequence (or upload a file for each).
  2. Choose the database (searches only):

    DatabaseContainsUse it to
    nr (non-redundant)Almost all protein sequences in NCBIRun the broadest search
    refseq_proteinCurated reference proteinsGet cleaner, less redundant hits
    swissprotManually reviewed proteins (UniProtKB/Swiss-Prot)Find well-characterised proteins
    pdbProteins with an experimental 3D structureFind structures to compare with your AlphaFold model

    These are blastp databases; blastn uses nucleotide databases such as core_nt or refseq_rna.

  3. Optional: limit the organism. Type a species or group in Organism (for example, Oryza sativa or Fungi) to restrict the search, or tick Exclude to remove it.

  4. Click BLAST. Searches take from seconds to a few minutes. Bookmark the results page or note the Request ID (RID): results are kept on the NCBI server for about 36 hours.

1.3Reading the resultsStep 3

The results page has four tabs:

TabWhat it shows
DescriptionsThe list of hits, best first, with their scores
Graphic SummaryWhere each hit aligns along your query, coloured by score
AlignmentsThe residue-by-residue alignment of your query with each hit
TaxonomyWhich organisms the hits come from

1.3.1Read the Descriptions table

Figure 1.3. The Descriptions table of a BLAST search against nr: hits are listed best first, with Max Score, Total Score, Query Cover, E value, Per. Ident, Acc. Len and Accession.
ColumnWhat it meansHow to read it
E valueNumber of hits this good expected by chanceThe smaller, the more significant. Below 1e-5 usually indicates real similarity; 0.0 means extremely significant
Query CoverPercentage of your query covered by the alignmentClose to 100% = similar over the whole length; low = only part matches, such as one domain
Per. IdentPercentage of identical residues in the aligned regionFor proteins, above 40% usually means closely related; 25–40% may still be related if E value and cover are good
Max ScoreScore of the best aligned regionHigher is better; used to rank hits
Total ScoreSum of scores of all aligned regionsHigher than Max Score when the hit aligns in several separate pieces
Acc. LenLength of the hit sequenceCompare with your query length
AccessionIdentifier of the hitClick it to open the full NCBI record

1.3.2Read the scores together

What you seeWhat it suggests
Low E value, high cover, high identityA close homolog, likely with the same function
Low E value, high cover, moderate identityA more distant homolog; function may be similar but should be checked
Low E value, low coverThe sequences share only a region, such as a common domain
High identity, very short alignmentOften a chance match; check the E value
E value above 0.01Probably no meaningful similarity

1.3.3Check the alignment

Check the alignment of your top hits in the Alignments tab (shortened example):

Query  1  MKCTNIFLTIGLVFLALLLGSQ  22
          MKC+NIFL IG+VFLALLLG+Q
Sbjct  1  MKCSNIFLAIGIVFLALLLGAQ  22
Symbol in the middle lineMeaning
A letterIdentical residue in both sequences
+Different but similar residues (proteins only)
BlankDifferent residues
- (in Query or Sbjct)A gap: a residue missing in one sequence
Figure 1.4. Alignments tab: the query is written above the subject (Sbjct). The first alignment is identical over its whole length (91/91 identities); the second is weak (15/87 identities, 37/87 positives), with + marking similar residues.

Above each alignment, Identities, Positives and Gaps summarise these counts.

1.4Troubleshooting

  • No significant hits: check that the program matches your sequence (DNA into blastn, protein into blastp), remove any organism limit, or try a broader database such as nr.
  • Error about the sequence: remove numbers, spaces or non-sequence characters, and keep only one header line starting with >.
  • Too many hits from the same species: use the Organism field to exclude it, or switch to refseq_protein or swissprot.
  • Top hits are all “hypothetical” or “uncharacterized” proteins: search swissprot to find characterised proteins.
  • The search is slow: shorter queries and smaller databases run faster; keep the page open or save the RID.

1.5References

  • Camacho C. et al. (2009). BLAST+: architecture and applications. BMC Bioinformatics 10: 421.
  • NCBI BLAST help and documentation
  • AI in Genomics 2026, Practical 3: Compare.

2Genome annotation

Prokka and RAST · 19 min read

Genome annotation predicts where the genes are in an assembly and assigns a function to each of them. This chapter describes three ways to annotate an assembly: Prokka on the command line (2.1), Prokka in your browser on the Galaxy EU server (2.2), and the RAST web server (2.3).

RouteWhere it runsSet-upTimeWhat you get
2.1 Prokka CLIYour computer, from a terminalInstall Prokka (conda is recommended)2–20 minutes for a typical genomeGFF, GenBank, FASTA and table files
2.2 Prokka GUIGalaxy EU server, in a browserNothing to install; a free account is optional but recommendedDepends on the server queueThe same files as the command-line version
2.3 RASTRAST server, in a browserAccount that is approved by a personA few hours to a dayAnnotation organised into subsystems, browsed in the SEED Viewer

2.1Prokka CLI

Command-line interface

Prokka annotates a bacterial, archaeal or viral assembly in 2–20 minutes on a laptop. It takes an assembly (FASTA or FNA) as input and predicts CDS, rRNA, tRNA, tmRNA and ncRNA features, then assigns products by matching against curated protein databases.

2.1.1Prokka installation and setup

Conda is the route that works on both Linux and macOS and pulls in all 20+ dependencies (BLAST+, Prodigal, Aragorn, Barrnap, HMMER, tbl2asn). Install Miniconda first if you do not have it.

# add the channels once, in this order
conda config --add channels defaults
conda config --add channels bioconda
conda config --add channels conda-forge

# install into its own environment (recommended)
conda create -n prokka_env -c conda-forge -c bioconda prokka
conda activate prokka_env

Alternatives if you do not use conda:

# macOS, Homebrew
brew install brewsci/bio/prokka

# Ubuntu / Debian
sudo apt-get install prokka

# Docker, no local dependencies at all
docker pull staphb/prokka:latest

Then index the databases and confirm the install:

prokka --setupdb      # builds the BLAST indices, run once after install
prokka --version      # should print e.g. prokka 1.14.6
prokka --listdb       # shows which kingdom and genus databases are available

If prokka --version fails with a Perl module error, the usual cause is a missing dependency outside conda; reinstalling inside a clean conda environment fixes it in almost every case. Remember to run conda activate prokka_env in every new terminal session.

2.1.2Annotation

Prokka takes the assembly as a positional argument: there is no input (-i) flag.

  1. Open a terminal.

    • On Ubuntu, Ctrl+Alt+T.
    • On macOS, “Terminal” from “Applications › Utilities”.
  2. Check where you are. This prints the current working directory, so you know what any relative path will be measured against.

    pwd
  3. Copy the path of the folder holding your assembly. In the file manager, right-click the folder and choose “Copy path”.

    • Ubuntu: Ctrl+L in Files shows the path.
    • macOS Finder: right-click and hold Option, then “Copy as Pathname”.
  4. Move into that folder. Type cd, a space, then paste. You will see something like this:

    cd /home/user/projects/genome_assembly

    OR drag and drop the folder onto the terminal window, which also pastes its path. Quote the path if it contains spaces.

  5. List the files and confirm the assembly is there using ls, which lists all the files in your directory.

    ls
  6. Run Prokka.

    prokka contigs.fasta --outdir results_prokka --prefix my_genome --compliant

What each part does:

  • contigs.fasta: your assembly file.
  • --outdir results_prokka: the folder to create for results. It must NOT already exist, or Prokka stops. If it does, add --force to your command to overwrite.
  • --prefix my_genome: names every output file my_genome.gff, my_genome.faa and so on. Without it they are named after today’s date, which gets confusing fast.
  • --compliant: enforces GenBank/ENA/DDBJ submission rules. It is shorthand for --addgenes --mincontiglen 200 --centre XXX, so it drops contigs under 200 bp and adds gene features alongside each CDS.
  • Optional: --cpus 4: threads to use; --cpus 0 uses every available core.

A typical bacterial genome finishes in 2–20 minutes. Progress is printed to the terminal and mirrored in the .log file.

2.1.3Optional arguments worth adding

Prokka accepts single or double dashes interchangeably (-compliant and --compliant both work). prokka --help lists everything.

ArgumentWhat it doesWhen to use it
--forceOverwrites an existing output directoryRe-running after a failed or tweaked attempt
--genus EscherichiaSets the genus in the annotation metadataAlways, if you know your organism
--species coliSets the speciesWith --genus, for cleaner headers
--strain K12Sets the strain nameMulti-isolate projects
--usegenusUses the genus-specific BLAST database instead of the generic oneMuch better product names, if your genus is in --listdb
--kingdom ArchaeaSelects the genetic code and databases: Bacteria (default), Archaea, Viruses, MitochondriaAnything that is not a bacterium
--gcode 4Overrides the translation tableMycoplasma, Spiroplasma and other code-4 organisms
--gram negRuns SignalP for signal-peptide predictionSecretome or surface-protein work
--locustag ECOPrefix for locus tags (e.g. ECO_00001)Submission, or keeping isolates distinct
--centre UoTSequencing centre ID written into the filesReplaces the placeholder XXX that --compliant inserts
--mincontiglen 500Skips contigs shorter than thisFragmented assemblies full of short junk contigs
--rfamAlso searches Rfam for ncRNAs with InfernalNon-coding RNA is of interest; roughly doubles runtime
--metagenomeTunes Prodigal for fragmented or mixed inputMAGs and metagenome bins
--proteins ref.gbkAnnotates against your own trusted proteins firstMatching a close reference strain’s nomenclature
--evalue 1e-09Tightens the similarity cutoff (default 1e-06)Reducing spurious product assignments
--norrna / --notrnaSkips rRNA or tRNA predictionSpeed, when only CDS matter
--quietSuppresses screen outputRunning inside a loop or pipeline

A fuller command for a known organism:

prokka contigs.fasta --outdir results_prokka --prefix ecoli_K12 \
       --genus Escherichia --species coli --strain K12 --usegenus \
       --locustag ECO --centre UoT --compliant \
       --cpus 8 --rfam

2.1.4Annotation outputs

Every file shares the --prefix you set, so with --prefix my_genome you get my_genome.gff, my_genome.faa and so on.

FileWhat it containsWhat you use it for
.gffGFF3: every predicted feature with coordinates, strand and product, plus the contig sequences at the end.The master annotation file. Feed it to Roary, Artemis, IGV or any pangenome tool.
.gbkThe same annotation in GenBank flat-file format, sequence and features together.Viewing in Artemis or SnapGene, and as input to tools that expect GenBank.
.faaAmino-acid FASTA of every translated CDS, headers carrying the locus tag and product.The workhorse for downstream function: BLAST, eggNOG, KEGG, InterProScan, AMR screening.
.ffnNucleotide FASTA of all transcripts: CDS plus rRNA, tRNA, tmRNA and misc_RNA.Gene-level nucleotide analyses, primer design, per-gene alignments.
.fnaNucleotide FASTA of the input contigs, unchanged apart from renamed headers.The reference copy of the assembly that matches the annotation coordinates.
.tsvTab-separated table: locus_tag, feature type, length, gene name, EC number, COG and product.Open in Excel or R to search, filter and count genes. Usually the first file to look at.
.txtA short statistics summary: contig count, bases, and counts of CDS, rRNA, tRNA, tmRNA.Quick sanity check: a 4 Mb genome with 300 CDS means something went wrong.
.tblNCBI feature table, the plain-text intermediate used to build the .sqn.Only touched during GenBank submission or manual feature correction.
.sqnASN.1 Sequin file generated by tbl2asn, bundling sequence and annotation.The file you actually submit to GenBank.
.fsaNucleotide FASTA of the contigs with extra Sequin tags in the headers.Input for the .sqn; ignore it unless you are submitting.
.errNCBI discrepancy report listing annotations that break submission rules.Read it before submitting; it tells you exactly what a curator would reject.
.logFull run log: every command executed and the version of each dependency.Reproducibility, methods sections, and debugging a run that looks odd.

Quick checks after a run

cd results_prokka
cat my_genome.txt                  # feature counts
head -3 my_genome.tsv              # table header and first rows
grep -c ">" my_genome.faa          # number of proteins predicted

As a rough guide, a typical bacterial genome yields around 900 CDS per Mb. Far fewer usually means a fragmented assembly or the wrong --kingdom.

2.1.5Resources and references

2.2Prokka GUI

On the Galaxy EU server

This workflow guides you through the annotation of your genome using the web-based version of Prokka, which is available on the Galaxy EU Server.

2.2.1Overview of Galaxy EU

The European Galaxy Server (usegalaxy.eu) is a public web platform for accessible, reproducible, and transparent computational research, primarily used for data analysis in fields like bioinformatics, genomics, and climate science.

Users can interact with thousands of specialized scientific tools via a graphical web interface rather than writing code.

2.2.2Running Prokka on the Galaxy EU Server

Set up your Galaxy EU account

  1. Open usegalaxy.eu in any browser.
  2. Click Login or Register in the top menu and create a free account. You can run tools anonymously, but an account keeps your histories, gives a larger storage quota and lets Galaxy email you when a job finishes.
  3. Start a new history for this analysis: in the “> History” panel on the right, click the + icon and rename it (click the name) to something like Prokka – isolate 12. A history is Galaxy’s project folder; every file you upload and every result appears there as a numbered dataset.

The screen has three panels: tools on the left, the working area in the centre where tool forms open, and your history on the right.

Search for Prokka

  1. In the search tools box at the top of the left panel, type Prokka.
  2. Click Prokka – Prokaryotic genome annotation. The tool form opens in the centre panel.
  3. Check the version button at the top of the form and use the newest one listed (1.14.6 builds at the time of writing), so results match the command-line version.

You can open the form now, but it only becomes useful once your genome is uploaded: the input dropdown lists datasets from your current history.

2.2.3Upload your genome

This is where most participants get stuck. The upload itself is easy; the trap is that Galaxy must recognise the file as fasta, or it will not appear in Prokka’s input list.

Upload the file

  1. Under Contigs to annotate, click on “...”. Click the Upload icon (an arrow pointing up) at the bottom left of the panel. The upload dialog opens.
  2. Add your assembly one of two ways:
    • Choose local file: pick the contigs FASTA from your computer.
    • Paste/Fetch data: paste a URL, for example an NCBI or Zenodo link, and Galaxy downloads it directly.
  3. Set the Type column to fasta. It defaults to auto-detect, which usually works but sometimes labels an assembly as txt. Setting it yourself removes the guesswork. Leave Genome as ?.
  4. Click Start, then Close the dialog. You do not need to keep it open.

Galaxy decompresses .gz and .zip files automatically, so a gzipped assembly is fine as it is. Upload the assembly, not the raw reads (.fastq): Prokka annotates contigs.

Check it arrived correctly and run the tool

The dataset appears at the top of your history and changes colour as it goes: grey queued, orange running, green done, red failed.

  1. When it is green, this means that the upload is done, NOT the annotation. Click its name to expand it and check the number of your upload (your first upload would be 1).
  2. Click the selection box under Contigs to annotate and choose the number that corresponds to your upload. In case you upload multiple genomes at once, this tells the tool which genome to annotate.
  3. Click Run Tool at the top right of the central panel.
  4. Output files will start showing in your history: grey (queued), orange (running), green (complete), or red (failed).
  5. Once complete, you may download the output file of your choice using the save icon that is found in each green box. OR click the eye icon to preview the contents; you should see > header lines followed by sequence.
Notes
  • When the outputs turn green they are named like Prokka on data 1: gff, Prokka on data 1: faa. The file contents are identical to the command-line version, which explains what each one holds (see the output table in 2.1 Prokka CLI).
  • One file at a time: click the dataset name to expand it, then the download icon (a floppy disk). Galaxy names the download after the dataset, so rename it on your computer, for example to isolate12.gff.
  • Everything at once: open the history options menu (the icon at the top of the history panel) and choose Export History to File. Galaxy packages every dataset into one archive you can download. This also keeps a record of the exact tool versions and settings, which is handy for methods sections.
  • Or keep working in Galaxy: the outputs are already in place for your next steps.
  • You can switch on Email notification at the bottom to receive updates on your jobs via email.

2.2.4Troubleshooting

ProblemWhat you seeFix
Wrong formatFormat shows txt or tabular; the file is missing from Prokka’s input dropdownClick the pencil (Edit attributes) → Datatypes tab → choose fasta → Save. No re-upload needed.
Uploaded reads insteadFormat shows fastqsangerProkka needs an assembly. Assemble first (for example with Shovill or SPAdes, both on Galaxy EU), or upload the contigs file instead.
Upload stays greyNothing happens for a long timeThe server queue is busy. Wait a few minutes before re-uploading; duplicates only add to the queue.
Upload turns redError on the datasetClick the bug icon to read the error. Usually a broken download link or an empty file.
Contig names too longProkka later fails with a message about contig ID lengthProkka rejects contig IDs over 37 characters, which long SPAdes names can exceed. Shorten the headers before uploading, then re-run.

2.2.5Sources

2.3RAST

Rapid Annotation using Subsystems Technology

The Rapid Annotation using Subsystems Technology (RAST; rast.nmpdr.org) annotates an uploaded assembly on a remote server and returns it organised into subsystems (curated functional groups of genes), which you browse in the SEED Viewer. There is nothing to install, but jobs are queued and typically take a few hours to a day.

2.3.1Create a RAST account

Accounts are approved by a human, so register a few days before you need results.

  1. Go to rast.nmpdr.org.
  2. Click Register for a RAST account in the left-hand menu.
  3. Fill in the form: name, institutional email, institution, and a short note on what you intend to annotate. An academic or institutional email address is approved faster than a personal one.
  4. Submit, then wait for the approval email. This usually arrives within a working day, occasionally longer.
  5. Log in from the RAST home page with the username and password you chose.

One account covers unlimited jobs, and your genomes stay private to you unless you explicitly make them public.

Worth knowing The RAST team now points new users towards BV-BRC (which absorbed PATRIC), where the same RASTtk annotation service runs alongside comparative-genomics tools. Classic RAST remains available and is what this guide describes, because the SEED Viewer subsystem browser is only there.

2.3.2Upload a new job

  1. Log in and click Upload New Job in the left menu.
  2. Choose the sequence file. Click Browse and select your assembly: a contigs FASTA (.fasta, .fna, .fa), optionally gzipped. Upload the assembly, not raw reads; RAST does not assemble.
  3. Click Use this data to move to the settings page.
  4. Fill in the organism taxonomic details.
    • Taxonomy ID: enter the NCBI taxid if you know it and click the lookup button; it auto-fills the rest. Otherwise leave it and type the fields manually.
    • Domain: Bacteria, Archaea or Virus.
    • Genus, species, strain: as precise as you can be. If unidentified, a genus plus “sp.” and your isolate code works.
    • Genetic code: 11 for most bacteria and archaea, 4 for Mycoplasma, Spiroplasma and relatives.
  5. Choose the annotation scheme. Select RASTtk, the current toolkit and the default. Classic RAST is kept only for reproducing old jobs.
  6. Set the options. The defaults are sensible; the ones worth a thought:
    • Automatically fix errors: leave ticked.
    • Fix frameshifts: tick it for draft assemblies, where sequencing errors create false frameshifts.
    • Build metabolic model: tick it if you want a draft flux-balance model in ModelSEED.
    • Backfill gaps: a second pass to catch genes missed on the first; harmless to leave on.
  7. Click Finish the upload. You get a job ID immediately, and an email when the job completes.

Runtime depends on the queue and on genome size. Check progress under Jobs Overview, which shows each stage with a coloured status per stage.

2.3.3Find and read the results

When the email arrives, go to Jobs Overview and click the job, or use My Jobs. Two buttons matter: View details, which opens the SEED Viewer, and Download, which exports the annotation as GenBank, GFF3, EMBL, amino-acid FASTA, nucleotide FASTA or an Excel spreadsheet.

Clicking View details lands you on the organism genome metrics: contig count, genome size, GC content, number of coding sequences, number of RNAs, and the proportion of genes placed in subsystems.

To browse the actual annotation, click Browse annotated genome in SEED Viewer.

Subsystem statistics: the two views of the same data

A subsystem is a curated set of genes that together carry out one biological process: a pathway, a complex, a transport system. The subsystem statistics panel shows them two ways, side by side.

ViewWhat it showsBest for
Circular diagram (pie chart)Each slice is a top-level category (amino acids, carbohydrates, virulence, and so on), sized by how many genes fall in it and colour-coded.Seeing the shape of the genome at a glance, and spotting an unusually large category. Click a slice to filter the table to it.
Subsystem tableA sortable hierarchy: category → subcategory → subsystem, with the gene count in each row.Actually finding genes. Anything you want to click through to starts here.

The two are linked: clicking a pie slice filters the table, and the table is where you drill down. Use the pie to orient yourself, the table to work.

Subsystem coverage is typically 40–60% of genes. The rest are real genes that are simply not in a curated subsystem yet, most of them hypothetical proteins. A low percentage is not a failed annotation.

2.3.4Choose a subsystem in the table

The table is a three-level hierarchy. You expand your way down to a named subsystem, and that subsystem gives you its gene list.

  1. Pick a category. The top-level rows are broad: Amino Acids and Derivatives, Carbohydrates, Cell Wall and Capsule, Virulence · Disease · Defense, Membrane Transport, and so on. Click the category name, or the matching slice in the circular diagram.
  2. Expand to a subcategory. Inside Carbohydrates, for example, you get Central carbohydrate metabolism, Fermentation, Monosaccharides, and others.
  3. Click the subsystem itself. Inside Central carbohydrate metabolism you might pick Glycolysis and Gluconeogenesis or TCA Cycle. The subsystem name is the clickable link.
  4. Read the subsystem page. It has two parts:
    • The functional roles table: each role in the pathway, and which of your genes fills it. Blank rows are informative: a missing role may mean an incomplete pathway, or a gene the annotation missed.
    • The gene list, with a feature ID for each, of the form fig|83333.1.peg.1234. peg stands for protein-encoding gene; the number is the gene’s index in your genome.
  5. Click a feature ID to open that single gene.

A practical shortcut for a workshop exercise: Virulence · Disease · Defense › Resistance to antibiotics and toxic compounds gives participants an immediately interesting gene list in any isolate.

2.3.5One gene: neighbourhood and sequences

Clicking a feature ID (fig|83333.1.peg.1234) opens the Annotation Overview page for that gene.

The genetic environment

Scroll to the Compare Regions panel. It draws your gene as a red arrow in the centre of its contig, with the genes flanking it on either side, and stacks the equivalent region from closely related genomes underneath.

How to read it:

  • Arrow direction is the strand, so you can see at a glance which neighbours are co-oriented and might sit in one operon.
  • Colour equals homology. Genes sharing a colour across rows are orthologues; your red gene’s colour group tracks it through every genome shown.
  • A conserved block of the same colours in the same order across many rows is a strong hint of a functional unit: an operon or a gene cluster.
  • A break in synteny (neighbours that differ from the related genomes) often marks a mobile element, a genomic island, or a horizontal transfer event.
  • Hover over any arrow for its function; click it to jump to that gene’s own page. The controls above let you widen the window and change how many genomes are compared.

The FASTA sequences

On the same page, the sequence section gives both forms of the gene:

SequenceWhereTypical use
DNA (nucleotide)The DNA sequence link or tab on the feature pagePrimer design, cloning, checking the start codon and the region around it
Protein (amino acid)The protein sequence link or tab, alongside itBLAST against NCBI (see 1. Sequence alignment), domain search in InterPro or Pfam, structure prediction (see 3. Protein modelling)

Both display as plain FASTA with the feature ID in the header, so you can copy-paste straight into BLAST. There are also links out to run BLAST directly, and to the gene’s entry in related databases.

For many genes at once, do not scrape this page: go back to the job page and use Download to export the whole annotation as amino-acid or nucleotide FASTA.

2.3.6Sources

  • RAST server: registration, job submission and the SEED Viewer
  • BV-BRC: the successor platform, running the same RASTtk annotation service

3Protein modelling

AlphaFold 3 Server · 4 figures · 9 min read

GetGenome supports the sequencing of your organism of interest, providing an assembly file (.fasta) and a protein sequence file (.faa) that is generated from the annotation of the genome. This workflow guides you through the steps to model your protein of interest (POI) using the .faa file and AlphaFold3 (AF3) Server.

3.1Overview

AlphaFold 3 Server is an advanced artificial intelligence model developed by Google DeepMind and Isomorphic Labs that predicts the 3D structures and interactions of life’s cellular molecules (proteins, DNA, RNA, chemical modifications, and small molecule ligands or ions).

3.2The input file

The protein sequence file (.faa) contains the protein sequences from your genome, with identifiers and, where possible, predicted functions. You do not need to identify the proteins yourself; you only need to find your protein of interest in the file. (A .faa file is one of the outputs of genome annotation; see 2. Genome annotation.)

3.2.1Open your protein sequence fileStep 1

Open the .faa file in a text editor (for example “Notepad”, “Sublime Text”). Do not use Microsoft Word or Microsoft Excel as they do not support such file formats.

3.2.2Locate your amino acid sequenceStep 2

AlphaFold 3 Server requires the primary amino acid sequence of your protein as the input. The protein file links each amino acid sequence to an identifier.

  1. Use Find (Ctrl+F or Cmd+F) and type your protein’s name or symbol. A matching line looks like this (shortened):

    contig_12 annotation gene 803242 803583 . - . ID=g04521;Name=AVR-Pik;product=AVR-Pik effector
    contig_12 annotation mRNA 803242 803583 . - . ID=g04521.t1;Parent=g04521

    Note the identifier of the mRNA line (here g04521.t1): the protein sequence is usually stored under this identifier.

  2. If the gene has several mRNA lines (for example g04521.t1 and g04521.t2), the gene has more than one predicted version of its protein. Use the primary one: the version marked as canonical or representative in the annotation if it is marked, otherwise the first (.t1), which is usually the longest.

  3. Copy the sequence. Copy the primary amino acid sequence, which will end before the next >:

    >g04521.t1
    MKCTNIFLTIGLVFLALLLGSQAEAETFLTIGLVFLALL
    >g03881.t1

Optional (recommended): Save the amino acid sequence as a new text file, for example my_protein.faa.

3.2.3Checkpoints

Before modelling, look at the sequence:

CheckGood signWarning sign
First letterMAnything else: the annotation may have missed the start
Stop charactersNo * or ., or only one at the very endA * or . in the middle: the gene prediction is probably wrong or the gene is broken
LengthSimilar to known versions of your protein in UniProtMuch shorter or longer: check the annotation or ask the GetGenome team

3.2.4Troubleshooting

  • Search finds nothing: gene names in the annotation may differ from the ones you know. Try the predicted function, or use the BLAST route:
If your protein is not named in the file, take its amino acid sequence and search it with blastp (see 1. Sequence alignment). The top hit gives you the identifier.
  • Several different genes match: your protein may belong to a family with several members in the genome. Use BLAST against a known version to pick the closest one.
  • Identifier not found in the proteins file: search for the gene identifier instead of the mRNA identifier (for example g04521 rather than g04521.t1); naming conventions vary.

3.3Input to AlphaFold 3 ServerStep 3

  1. Sign in at alphafoldserver.com with your Google account.
  2. In the first entity box, choose Protein as the molecule type.
  3. Paste your sequence without the header line (the line starting with >). Remove any trailing * or ..
  4. To model your protein together with other molecules, click Add entity and add each one: another protein, DNA, RNA, a ligand or an ion.
  5. Give the job a descriptive name.
  6. Leave the seed* on automatic for a first run.
  7. Preview the job, check that every entity and copy number is correct, then Confirm and submit job.

The server limits the total size of each job and the number of jobs per day. Both are shown on the server.

* A seed is the random starting number for a prediction: the same seed reproduces the same result, and different seeds give alternative predictions whose agreement tells you how reliable the structure is.

3.3.1Number of copies

Each entity has a Copies field. It sets how many copies of that same molecule are in the model.

Set more than one copy only when there is evidence that the protein works that way, from the literature or from known structures of related proteins. AlphaFold will build a complex from whatever copies you give it, and the result can look plausible even when that complex does not exist in nature.

3.4Understand the outputStep 4

When the job finishes, open it from the job list. The results page shows:

  • the 3D model, coloured by confidence (see below);
  • the pTM and, for complexes, ipTM scores;
  • the green plot (see below);
  • a selector to switch between the five models AlphaFold produces for each job.

Click Download to get a zip file:

FileContent
…_model_0.cif to …_model_4.cifThe five predicted structures. model_0 is the top-ranked one
…_summary_confidences_0.json to _4.jsonThe overall scores of each model (pTM, ipTM, ranking score and others)
…_full_data_0.json to _4.jsonPer-residue and per-pair confidence values, including the data behind the green plot
…_job_request.jsonExactly what you submitted. Keep it so the job can be reproduced

The .cif files open in UCSF ChimeraX (see 4. Protein visualization).

3.5Interpreting the confidence scores

Read the confidence scores before looking at any structural detail. A model without its confidence scores should not be interpreted.

3.5.1Colours on the model (pLDDT)

The model is coloured residue by residue by how confident AlphaFold is in the local structure:

ColourScoreHow to read it
Dark blue> 90Very high confidence: backbone and side chains likely accurate
Light blue70–90Confident: the backbone is likely correct
Yellow50–70Low confidence: treat with caution
Orange< 50Very low confidence: do not interpret the shape of this region

Orange regions are often flexible or disordered parts of the protein. AlphaFold still draws them, sometimes as long loops or isolated helices, but their shape carries no information.

Figure 3.1. Three AlphaFold models coloured by pLDDT, with the colour key. The right-hand model is dark blue (very high confidence) throughout, while the left-hand model has many low-confidence yellow and orange regions. Source: EMBL-EBI online tutorial, AlphaFold: a practical guide.

3.5.2pTM

pTM is a single score from 0 to 1 for the overall fold of the whole model.

pTMHow to read it
> 0.5The overall fold may be correct
< 0.5The overall fold is likely wrong or uncertain

3.5.3ipTM

ipTM appears only when the job has more than one entity or copy. It scores, from 0 to 1, how confident AlphaFold is in the way the entities are placed relative to each other.

ipTMHow to read it
> 0.8The interaction between the entities is confidently predicted.
0.6–0.8Grey zone: the interaction may be correctly or falsely predicted. Treat cautiously.
< 0.6The interaction is likely wrongly predicted.

pTM and ipTM are less reliable for very small proteins and short chains; for these, rely more on the colours and the green plot.

Figure 3.2. pTM and ipTM when the Pikp-HMA domain is modelled as a monomer, dimer, trimer and tetramer. The dimer has the highest ipTM (0.83). Black boxes mark the individual copies along the diagonal.

3.5.4The PAE plot (Predicted Aligned Error)

The PAE plot shows how confident AlphaFold is in the relative position of every part of the model against every other part.

  • Both axes list the residues of the model, in order, chain after chain. When there are several chains, lines on the plot mark where one chain ends and the next begins.
  • Each point shows the expected error in the position of the residue on one axis when the model is aligned on the residue on the other axis.
  • Dark green means low expected error (confident); light or white means high expected error (not confident).
Figure 3.3. PAE plot of Pikp-1 (residues 186–486). The HMA and NB-ARC domains are dark green squares on the diagonal; the lighter green between them means their position relative to each other is less certain.

How to read common patterns:

PatternMeaning
Dark square along the diagonalA confidently modelled domain
Two dark squares on the diagonal, light between themTwo domains each modelled confidently, but their position relative to each other is uncertain
Dark blocks off the diagonal, between two chainsThe two chains are confidently placed relative to each other; supports the predicted interface
Light blocks between two chainsAlphaFold is not confident how the chains are arranged, whatever the model looks like
Figure 3.4. Schematic PAE plots for two chains, A and B, with the pTM and ipTM each pattern usually gives: (1) low confidence overall; (2) each chain confident, but not their relative position; (3) confident only at the interface; (4) full relative confidence.

The colours on the model and the green plot answer different questions. A region can be dark blue on the model, meaning its local structure is confident, while the green plot shows its position relative to another domain or chain is uncertain.

3.5.5Reading the scores together

What you seeWhat it suggests
Mostly blue model, pTM > 0.5The fold is likely reliable
Blue chains, high ipTM, dark green between chainsThe complex is confidently predicted
Blue chains, low ipTM, light green between chainsEach chain is folded confidently, but the interaction is not supported
Mostly orange modelPossible disordered protein, missing partner, or wrong sequence (go back to Step 2)
The five models look very differentThe prediction is uncertain, even if one model scores well

A confident model is a well-founded hypothesis about your protein’s structure, not proof. Use it to decide which experiments are worth doing.

3.6Troubleshooting

  • “Invalid sequence”: a header line, a * or ., spaces, or letters that are not amino acids were pasted.
  • Job too large: model a single domain, or reduce the number of copies or partners.
  • Daily limit reached: the counter resets; submit the job the next day.
  • Everything is orange: check the sequence (Step 2 checkpoints) before re-running.

3.7References

  • Abramson J. et al. (2024). Accurate structure prediction of biomolecular interactions with AlphaFold 3. Nature.
  • Maqbool A. et al. (2015). Structural basis of pathogen recognition by an integrated HMA domain in a plant NLR immune receptor. PMID 26304198.
  • Posbeyikian A., Toghani A., McClelland, Kamoun S., Sugihara Y., Contreras (2026). Fold on Tight: Ten Things to Know About AI Structure Prediction in Plant Immunity.

4Protein visualization

UCSF ChimeraX · 8 figures · 6 min read

UCSF ChimeraX is a free program to visualise and analyse 3D protein structures. This workflow guides you through installing ChimeraX and learning its basic commands.

4.1InstallationStep 1

Download ChimeraX from the official download page, choosing the version for your operating system. It is free for academic and non-commercial use: accept the licence to start the download. Install the latest version, as AlphaFold 3 files need a recent one.

Operating systemHow to install
WindowsRun the .exe installer and follow the prompts.
macOSOpen the .dmg file and drag ChimeraX into Applications. Apple silicon (M1 and later) and Intel Macs have separate downloads.
LinuxInstall the .deb (Ubuntu, Debian) or .rpm (Fedora, Red Hat) package.

Open ChimeraX and type version in the command line at the bottom of the window: the version number appears in the Log.

4.1.1Troubleshooting

  • Installation is blocked: institutional laptops often need administrator rights; ask your IT department.
  • macOS says the app cannot be opened: go to System Settings → Privacy & Security and click Open Anyway.

4.2Using ChimeraXStep 2

ChimeraX is used through menus and through commands typed in the command line. Every command and its result appear in the Log.

Part of the windowWhat it does
Command line (bottom)Type a command and press Enter.
LogShows commands, results and error messages.
Toolbar tabs (Home, Molecule Display…)One-click actions, such as colouring by chain or taking a picture.
Models panelLists open models; show, hide or close them.
Sequence iconShows the sequence of each chain, linked to the 3D structure.

4.2.1Open a structure

Open a structure from the Protein Data Bank (PDB) with its four-character ID:

open 7B1I

Click [more info...] in the Log to see its metadata, such as the paper title and the experimental method.

4.2.2Point to parts of a structure

Use these symbols, combined from left to right (for example #1/b:1-10):

SymbolLevelExample
#Model#1 = the first model opened
/Chain/A = chain A
:Residue:51 = residue 51; :1-10 = residues 1 to 10; :1,5,20 = residues 1, 5 and 20
@Atom@ca = alpha carbon atoms

4.2.3Select, colour and display

CommandWhat it does
sel #1/cSelects chain C
color sel grayColours the selection gray
color #1/c #2B89ADColours chain C with a hex colour code
color #1/b:3-9 blueColours residues 3 to 9 of chain B
color #1 bychainGives each chain its own colour
surface #1/cShows the surface of chain C
hide #1 surfacesHides all surfaces
show #1/a:51 atomsShows the atoms (side chain) of residue 51 of chain A
viewCentres the structure in the window
help colorOpens the documentation for any command
Figure 4.1. Structure 7B1I in ChimeraX with chain B coloured by secondary structure: α-helices in gray and β-sheets in orange.
Figure 4.2. A surface drawn with surface #1/c: AVR-Pik (surface) bound to OsHIPP19 (ribbon).

Close the structure before the next step:

close #1

4.3Open your AlphaFold 3 modelStep 3

Unzip your AlphaFold 3 download (see 3. Protein modelling) and drag …_model_0.cif (the top-ranked model) into the ChimeraX window, or type its path:

open ~/Downloads/fold_my_job/fold_my_job_model_0.cif
Figure 4.3. An AlphaFold 3 model (model_0.cif) opened in ChimeraX and coloured by chain with color #1 bychain.

4.3.1Colour by pLDDT

Blue = confident, orange = low confidence:

color bfactor #1 palette alphafold
Figure 4.4. The model coloured by pLDDT, with the colour key from high confidence (blue) through low confidence (yellow) to disordered regions (orange).

4.3.2Load the PAE plot

Go to Tools → Structure Prediction → AlphaFold Error Plot and select your model. ChimeraX usually finds …_full_data_0.json; if not, click Browse, select it, then Open. The same, by command:

alphafold pae #1 file ~/Downloads/fold_my_job/fold_my_job_full_data_0.json
Figure 4.5. The AlphaFold Error Plot window (top) and the PAE data it shows once the …_full_data_0.json file is loaded (bottom).

The PAE plot is interactive: drag a box over a region of the plot and it is highlighted on the 3D model. The plot window also recolours the model by pLDDT or by PAE domains.

Figure 4.6. The PAE plot is interactive: the region selected on the plot is highlighted on the 3D model (here in pink and green).

4.4Analyse the interfaceStep 4

4.4.1Show contacts

Show contacts between chain A and chain B closer than 3.5 Å. Each contact is drawn as a line (pseudobond) coloured by PAE, from blue (confident) to red (not confident). The PAE plot must be loaded first (Step 3).

alphafold contacts #1/a to #1/b distance 3.5

To look at one residue only, here residue 20 of chain B within 4.5 Å of chain A:

alphafold contacts #1/b:20 to #1/a distance 4.5
Figure 4.7. Contacts between two chains drawn as pseudobonds coloured by PAE: blue (0) is confident, red (20) is poor.

4.4.2Export the contacts

Export the contacts to a text file, to sort them in Excel:

alphafold contacts #1/a to #1/b distance 3.5 outputFile AF3-contacts.txt

Each row is a pair of residues in contact and its PAE, for example /A:81 /B:54 17.50.

PAE (Å)Confidence
0–10Confident
10–15Grey zone
> 15Poor
Figure 4.8. Contacts exported to AF3-contacts.txt (left) and sorted by PAE in Excel (right): low values (green) are confident contacts, high values (orange) are poor.

Residues in confident contacts are candidates to mutate when testing the interaction in the lab.

4.5Save your workStep 5

ChimeraX does not save automatically.

CommandWhat it saves
save ~/Desktop/my_session.cxsThe whole session; reopen it with open
save ~/Desktop/image.png supersample 3A high-quality picture of the current view

To record a 360° spin movie, run these four commands in order:

movie record
turn y 2 180
wait 180
movie encode ~/Desktop/spin.mp4

Pictures and movies can also be taken from Home → Images. Type pwd to see the folder where files are saved. On Windows, you can also write full paths, such as C:\Users\yourname\Desktop\image.png.

4.6Troubleshooting

  • Command not recognised: check the spelling, or type help followed by the command name.
  • Wrong residues coloured: check the model number (#), chain ID (/) and residue numbers in the Models panel and the Sequence viewer.
  • The PAE plot does not load: select the full_data file with the same number as the model (model_0 with full_data_0).
  • The contacts command gives an error: load the PAE plot first (Step 3).
  • File not found: the path is wrong; drag the file into ChimeraX instead.

4.7References

  • Meng E.C. et al. (2023). UCSF ChimeraX: Tools for structure building and analysis. Protein Science 32: e4792.
  • ChimeraX User Guide
  • AI in Genomics 2026, Practical 2: Interpret (slides).

5Structure alignment

Foldseek and FoldMason · 2 figures · 5 min read

Foldseek compares a protein structure with millions of known and predicted structures and finds the most similar ones, even when their sequences are unrelated. This workflow guides you through preparing your structure, running a search, reading the results, and aligning several structures with FoldMason.

5.1Before you startStep 1

Get everything you can from the sequence first: run BLAST (see 1. Sequence alignment) and InterProScan. Use Foldseek when the sequence gives few or no answers.

Foldseek needs a structure file in PDB or CIF format, either an experimental structure or a prediction (for example, …_model_0.cif from AlphaFold 3; see 3. Protein modelling). Your results are only as good as the structure you put in, so check it first:

CheckWhyHow
ConfidenceLow-confidence regions give misleading matchesColour by pLDDT in ChimeraX (see 4. Protein visualization); remove regions below 50 if they are long
One chainA complex mixes the signals of several proteinsKeep only your protein’s chain
FormatFoldseek reads PDB and mmCIF filesConvert in ChimeraX if needed

To keep one chain and save it as a PDB file in ChimeraX (here, removing chain B):

open model_0.cif
delete /B
save my_protein.pdb

5.2Searching with FoldseekStep 2

Go to Foldseek Search.

Figure 5.1. Foldseek search page: upload the structure, choose the mode, and click Search.
  1. Upload your structure with UPLOAD PDB, or use LOAD ACCESSION to load a structure by its identifier.
  2. Choose the databases. Keep all of them for the broadest search, or untick the ones you do not need:

    DatabaseContains
    PDB100Experimental structures from the Protein Data Bank
    AlphaFold/Swiss-ProtPredicted structures of manually reviewed proteins
    AlphaFold/UniProt50Predicted structures covering most known proteins
    AlphaFold/ProteomePredicted structures from selected model organisms
    CATH50Classified protein domains
    Other databasesPredicted structures from metagenomes, viruses and other sources
  3. Choose the mode:

    ModeComparesUse it to
    3Di/AALocal similarity, using structure and sequence togetherRun a fast, sensitive search for related proteins (default)
    TM-alignOverall (global) structural similarityTest whether hits share the whole fold
    LoL-alignLocal structural similarityFind shared regions that suggest a common origin
  4. Optional settings. Use Taxonomic filter to restrict hits to a group of organisms. Tick Iterative search to find more distant relatives.
  5. Click SEARCH. Results usually appear within a few minutes. Bookmark the page, or download the results to reload them later with UPLOAD PREVIOUS RESULTS.

5.3Reading the Foldseek resultsStep 3

Results are listed per database, best hits first. Click a hit to see its alignment with your query and the two structures superposed.

ColumnWhat it meansHow to read it
TM-scoreOverall structural similarity, from 0 to 1 (TM-align mode)Below 0.2: unrelated; above 0.5: generally the same fold
Prob.Probability that the hit is a true homologClose to 1: likely related
E-ValueNumber of hits this good expected by chanceThe smaller, the more significant
Seq. Id.Percentage of identical residues in the alignmentLow identity with a high TM-score suggests a hidden relationship
Pos. in Query / TargetWhere the alignment starts and ends in each proteinShows whether the whole protein or only part of it matches

5.3.1Read the scores together

What you seeWhat it suggests
High TM-score, high sequence identityA close homolog; BLAST should also find it
High TM-score, low sequence identitySame fold without detectable sequence similarity
High score over only part of the queryThe proteins share a domain or region
TM-score between 0.2 and 0.5Uncertain; check the superposition and other tools

A similar structure suggests, but does not prove, a shared origin or function. No single tool gives the whole answer: combine Foldseek with BLAST, domain annotation and the literature.

5.4Aligning many structures with FoldMasonStep 4

To compare several structures at once, for example your query and its best hits, use FoldMason, on the same website.

Figure 5.2. FoldMason page: upload at least two structures and click Align.
  1. Upload at least two structures (PDB or mmCIF): click CLICK TO SELECT FILES or drag and drop them.
  2. Click ALIGN.

The result is a structure-based multiple alignment: residues that occupy the same position in 3D are aligned, even when the sequences differ. Look for columns that are conserved across all structures, such as cysteines or glycines, which often hold the fold together.

5.5Troubleshooting

  • The upload is rejected: check the file is PDB or mmCIF; re-save it from ChimeraX.
  • Hits match a partner protein instead of yours: the file contains several chains; keep only your protein (Step 1).
  • Many weak hits: switch to TM-align mode and focus on hits with TM-score above 0.5.
  • No hits: try 3Di/AA mode with iterative search, and keep all databases selected.
  • Hits only match a disordered or low-confidence region: remove low-pLDDT regions and search again.

5.6References

  • van Kempen M. et al. (2024). Fast and accurate protein structure search with Foldseek. Nature Biotechnology 42: 243–246.
  • Xu J. and Zhang Y. (2010). How significant is a protein structure similarity with TM-score = 0.5? Bioinformatics 26: 889–895.
  • Foldseek on GitHub
  • AI in Genomics 2026, Practical 3: Compare (slides).

6Glossary

Key terms · 119 terms in 11 topics

Key terms used across the AI in Genomics workflows, grouped by topic and listed from A to Z within each topic.

Filter by topic

Genomics fundamentals13 terms

Accessory genome
Genes present in only some individuals or strains of a species; often linked to adaptation, such as virulence or resistance.
Chromosome
A long DNA molecule carrying many genes, packaged with proteins; each species has a typical number of chromosomes.
Contig
A continuous stretch of sequence assembled from overlapping reads, with no gaps.
Core genome
Genes shared by all individuals or strains of a species; they usually carry essential functions.
GC content
The percentage of G and C bases in a DNA sequence; it varies between species and between genome regions.
Genome
The complete DNA of an organism, including all its genes and non-coding sequences.
Genome size
The total length of an organism's genome, usually given in base pairs (bp, Mb or Gb).
Haplotype
A set of variants inherited together on the same copy of a chromosome.
Pangenome
The full set of genes found across all individuals of a species: the core genome plus the accessory genome.
Ploidy
The number of complete chromosome sets in a cell: haploid (one), diploid (two) or polyploid (more than two).
Reference genome
A representative genome assembly of a species, used as a standard to map reads and compare other genomes.
Scaffold
Contigs ordered and oriented relative to each other, with gaps of estimated size between them.
Synteny
Conservation of gene order along chromosomes between species or genomes.

Genome assembly & quality9 terms

Completeness
How much of the expected gene content is present in an assembly, often estimated with BUSCO.
Contamination
Sequences in an assembly that come from another organism, such as bacteria or the host.
Coverage
The fraction of the genome covered by at least one read; also often used to mean average depth.
De novo assembly
Building a genome sequence from reads without using a reference genome.
Depth
The average number of reads covering each position of the genome, for example 30×.
k-mer
A subsequence of length k; k-mer counts are used to assemble genomes and estimate genome size.
L50
The smallest number of contigs that together contain half of the assembly length; lower is better.
N50
The length of the shortest contig among the longest contigs that together make up half of the assembly; higher is better.
Reads
DNA sequences produced by a sequencing machine, short or long, which are assembled or mapped to a reference.

Genome annotation11 terms

CDS (coding sequence)
The part of a gene that is translated into protein, from the start codon to the stop codon.
Exon
A part of a gene kept in the mature mRNA after splicing; it can include coding and untranslated sequence.
Functional annotation
Assigning predicted functions to genes, such as names, domains or GO terms, usually by similarity to known proteins.
Gene model
The predicted structure of a gene: its position, exons, introns, CDS and UTRs.
Gene prediction
Identifying the position and structure of genes in a genome with software, experimental evidence, or both.
Intron
A non-coding part of a gene, removed from the mRNA by splicing.
ORF (open reading frame)
A stretch of DNA from a start codon to a stop codon, with no stop codon in between, that may encode a protein.
Pseudogene
A gene copy that has lost its function, often through mutations such as premature stop codons.
Structural annotation
Identifying where genes and their parts (exons, introns, CDS, UTRs) are located in a genome.
Transcript/protein prediction
Deriving the mRNA and protein sequences of each gene from its gene model.
UTR (untranslated region)
The parts of an mRNA before (5′ UTR) and after (3′ UTR) the coding sequence, which are not translated.

Sequence analysis13 terms

Alignment
Arranging two or more sequences so that corresponding residues line up, with gaps where needed.
BLAST
A tool that searches databases for sequences similar to a query and scores how significant each match is.
Conservation
How unchanged a position or region remains across related sequences; conserved regions are often important for function.
Domain
A region of a protein that folds and functions largely on its own and is found in different proteins.
Global alignment
An alignment over the full length of both sequences.
HMM (hidden Markov model)
A statistical model that describes a sequence as a series of hidden states; widely used to detect patterns in sequences.
Local alignment
An alignment of only the most similar regions of two sequences; the approach used by BLAST.
Motif
A short, recurring sequence pattern, often with a specific function, such as a binding site.
Multiple sequence alignment (MSA)
An alignment of three or more sequences, used to find conserved positions and build phylogenetic trees.
Profile HMM
An HMM built from a multiple sequence alignment, describing which residues each position of a family allows; used by Pfam.
Sequence identity
The percentage of aligned positions that have identical residues.
Sequence similarity
The percentage of aligned positions that have identical or chemically similar residues.
Whole-genome alignment
Aligning entire genomes to find conserved regions, rearrangements and variants.

Protein structure & structural bioinformatics6 terms

ipTM
AlphaFold score (0 to 1) of confidence in the interfaces between entities; above 0.8 indicates a confident complex.
PAE (predicted aligned error)
AlphaFold estimate, in Å, of the error in the relative position of two residues; low values mean confident placement.
pLDDT
AlphaFold per-residue confidence score (0 to 100) in the local structure; above 70 is generally reliable.
pTM
AlphaFold score (0 to 1) of confidence in the overall fold; above 0.5 suggests the fold may be correct.
RMSD (root mean square deviation)
The average distance, in Å, between matching atoms of two superposed structures; lower means more similar.
TM-score
A score (0 to 1) of the overall similarity of two structures; above 0.5 generally indicates the same fold.

Population genomics9 terms

Admixture
Mixing of genetic ancestry from previously separated populations; also the name of a tool that estimates it.
FST
A measure (0 to 1) of genetic differentiation between populations: 0 means none, 1 means completely separated.
Genetic diversity
The amount of genetic variation within a population or species.
GWAS (genome-wide association study)
A study that links genetic variants across the genome to a trait or disease.
Linkage disequilibrium (LD)
The non-random association of alleles at different positions, often because they are close on a chromosome.
Nucleotide diversity (π)
The average number of differences per site between two sequences taken from a population.
PCA (principal component analysis)
A method that summarises many variables into a few axes; used to visualise population structure.
Population structure
Genetic differences between groups within a species, caused by limited mixing between them.
Selective sweep
Reduced diversity around a beneficial variant that rose quickly in frequency under selection.

Genomic data formats16 terms

BAM
The compressed, binary version of SAM; smaller and faster to read.
CRAM
A highly compressed alignment format that stores reads relative to a reference genome.
CSV (comma-separated values)
A table stored as plain text, with columns separated by commas.
FASTA
A text format for sequences: a header line starting with > followed by the DNA or protein sequence.
FASTQ
A text format for sequencing reads, storing each read's sequence and the quality score of each base.
GFF (General Feature Format)
A tab-separated format describing genome features, such as genes and exons, and their positions.
GTF (Gene Transfer Format)
A format similar to GFF, commonly used for gene and transcript annotations.
JSON
A text format that stores structured data as key–value pairs; AlphaFold confidence scores are saved as JSON.
Metadata
Information that describes data, such as sample origin, sequencing method or date.
mmCIF
The current standard format for 3D structures of molecules, replacing PDB format; AlphaFold 3 models are .cif files.
Newick
A text format that represents phylogenetic trees with nested parentheses, branch lengths and labels.
PDB format
An older text format for the 3D atomic coordinates of molecular structures.
SAM (Sequence Alignment/Map)
A text format that stores reads aligned to a reference genome.
TSV (tab-separated values)
A table stored as plain text, with columns separated by tabs.
VCF (Variant Call Format)
A text format listing genetic variants, such as SNPs and indels, and the genotype of each sample.
YAML
A human-readable text format for structured data, often used for configuration files.

Command line / HPC16 terms

Bash
The most common shell on Linux and macOS, and a scripting language to automate commands.
CPU
The processor that runs most computations; cluster jobs request a number of CPU cores.
Environment
An isolated set of software and versions, such as a conda environment, used to run analyses reproducibly.
GPU
A processor specialised in parallel computation; needed to run AI models such as AlphaFold efficiently.
Job
A task submitted to a computing cluster, which runs when the requested resources are available.
Linux
A free, open-source, Unix-like operating system used on most servers and computing clusters.
Modules
A system on computing clusters to load and unload pre-installed software, for example module load blast.
RAM
Working memory used during computation; large genomes and AI models need more RAM.
Resources
The CPUs, GPUs, memory and time a job requests from a computing cluster.
Scratch
Fast, temporary storage on a cluster for running jobs; files are often deleted automatically after a set time.
Shell
A program that reads commands typed in the terminal and runs them.
Slurm
A job scheduler that queues and manages jobs on computing clusters.
SSH (Secure Shell)
A secure way to log in to a remote computer, such as a cluster, from your terminal.
Storage
Long-term disk space for data and results, usually backed up, unlike scratch.
Terminal
A window where you type text commands to control a computer.
Unix
A family of operating systems on which Linux and macOS are based.

Programming & data analysis12 terms

API (application programming interface)
A defined way for programs to request data or services from another program or database, such as NCBI.
Biopython
A Python library for biological computation, such as reading sequence files and querying databases.
Data frame
A table of data with named columns, each holding one type of value; central to pandas and R.
Git
A version control system that records changes to files and lets you return to earlier versions.
GitHub
A website that hosts Git repositories, for sharing and collaborating on code.
Jupyter
Notebooks that combine code, results, plots and notes in one document, run in a web browser.
pandas
A Python library for working with tables (data frames).
Plotting
Showing data as graphs or charts, for example with matplotlib in Python or ggplot2 in R.
Python
A general-purpose programming language widely used in bioinformatics and machine learning.
R
A programming language for statistics and data visualisation, widely used in biology.
Regular expression
A pattern used to search or match text, for example to find all sequence identifiers in a file.
tidyverse
A collection of R packages for cleaning, transforming and plotting data, including dplyr and ggplot2.

AI/ML in genomics5 terms

Deep learning
Machine learning that uses neural networks with many layers; the basis of tools such as AlphaFold.
Generative adversarial network (GAN)
Two neural networks trained against each other: one generates data, the other judges whether it is real.
Machine learning
Methods that let computers learn patterns from data to make predictions, without being explicitly programmed.
Natural language processing (NLP)
AI methods for understanding and generating text; similar methods are used to model DNA and protein sequences.
Reinforcement learning
Machine learning in which a model learns by trial and error, guided by rewards.

Databases & resources9 terms

ENA (European Nucleotide Archive)
The European database of nucleotide sequences and sequencing reads, run by EMBL-EBI.
GenBank
NCBI's public database of nucleotide sequences submitted by researchers.
InterPro
A database that classifies proteins into families and predicts their domains by combining several databases, including Pfam.
NCBI
The US National Center for Biotechnology Information, which hosts major databases and tools such as GenBank and BLAST.
PDB (Protein Data Bank)
The archive of experimentally determined 3D structures of proteins and nucleic acids.
Pfam
A database of protein families and domains, each described by a profile HMM; now part of InterPro.
RefSeq
NCBI's curated, non-redundant set of reference sequences for genomes, transcripts and proteins.
SRA (Sequence Read Archive)
NCBI's database of raw sequencing reads.
UniProt
The main database of protein sequences and functions; its Swiss-Prot section is manually reviewed.