Contents
This wiki collects the GetGenome step-by-step workflows used in the AI in Genomics training: aligning a sequence, annotating a genome, modelling a protein, visualising the model and comparing structures. Each chapter can be read on its own, and the glossary at the end explains the terms used throughout.
- 1
Compare a DNA or protein sequence with millions of known sequences using NCBI BLAST, and read the results.
4 figures · 5 min read
- 2
Predict genes and assign functions to an assembly with Prokka on the command line (2.1), Prokka on Galaxy EU (2.2) or RAST (2.3).
19 min read
- 3
Model your protein of interest with the AlphaFold 3 Server and judge how far the model can be trusted.
4 figures · 9 min read
- 4
Install UCSF ChimeraX, open an AlphaFold 3 model and analyse its confidence and its interface.
8 figures · 6 min read
- 5
Search for similar structures with Foldseek and align several structures with FoldMason.
2 figures · 5 min read
- 6
Key terms used across the workflows, filterable by topic and listed from A to Z.
119 terms in 11 topics
1Sequence alignment
NCBI BLAST compares a DNA or protein sequence with millions of known sequences in the NCBI database and finds the most similar ones. This workflow guides you through choosing the right type of BLAST, running a search or an alignment, and reading the results.
1.1Choosing BLASTStep 1
Go to NCBI BLAST. The type of BLAST depends on what your sequence is and what you want to compare it with:
| Program | Your sequence | Compared with | Use it to |
|---|---|---|---|
| blastn | DNA | DNA | Find similar genes, or check a PCR product or clone |
| blastp | Protein | Protein | Find sequence-similar proteins and predict function |
| blastx | DNA, translated | Protein | Find which protein a DNA sequence may encode |
| tblastn | Protein | DNA, translated | Find genes encoding your protein in genomes |
| tblastx | DNA, translated | DNA, translated | Compare distantly related DNA sequences |
Translated programs read DNA in all six reading frames, so they find proteins even when you do not know the gene structure.
1.2Running BLASTStep 2
BLAST can be run in two ways:
| Search a database | Align two or more sequences | |
|---|---|---|
| Question | Which known sequences are similar to mine? | How similar are these specific sequences? |
| Compared with | A whole database, such as all proteins in NCBI | Only the sequences you paste or upload |
| How | Default setting | Tick Align two or more sequences below the query box |
| Example | Identify an unknown protein from your genome | Compare your protein with a known version from UniProt |
To align many sequences together (for example, for a phylogenetic tree), use a multiple alignment tool such as Clustal Omega or MAFFT instead: BLAST compares your query with each sequence separately.
Enter your query. Paste your sequence (FASTA format or bare sequence) or an NCBI accession number, or upload a file.
Figure 1.1. NCBI BLAST (blastp) search page for a database search: enter the query, leave “Align two or more sequences” unticked, choose the database and click BLAST. Figure 1.2. The same page set up to align your own sequences: tick “Align two or more sequences”, then enter the query and the subject sequence (or upload a file for each). Choose the database (searches only):
Database Contains Use it to nr (non-redundant) Almost all protein sequences in NCBI Run the broadest search refseq_protein Curated reference proteins Get cleaner, less redundant hits swissprot Manually reviewed proteins (UniProtKB/Swiss-Prot) Find well-characterised proteins pdb Proteins with an experimental 3D structure Find structures to compare with your AlphaFold model These are blastp databases; blastn uses nucleotide databases such as core_nt or refseq_rna.
Optional: limit the organism. Type a species or group in Organism (for example, Oryza sativa or Fungi) to restrict the search, or tick Exclude to remove it.
Click BLAST. Searches take from seconds to a few minutes. Bookmark the results page or note the Request ID (RID): results are kept on the NCBI server for about 36 hours.
1.3Reading the resultsStep 3
The results page has four tabs:
| Tab | What it shows |
|---|---|
| Descriptions | The list of hits, best first, with their scores |
| Graphic Summary | Where each hit aligns along your query, coloured by score |
| Alignments | The residue-by-residue alignment of your query with each hit |
| Taxonomy | Which organisms the hits come from |
1.3.1Read the Descriptions table
| Column | What it means | How to read it |
|---|---|---|
| E value | Number of hits this good expected by chance | The smaller, the more significant. Below 1e-5 usually indicates real similarity; 0.0 means extremely significant |
| Query Cover | Percentage of your query covered by the alignment | Close to 100% = similar over the whole length; low = only part matches, such as one domain |
| Per. Ident | Percentage of identical residues in the aligned region | For proteins, above 40% usually means closely related; 25–40% may still be related if E value and cover are good |
| Max Score | Score of the best aligned region | Higher is better; used to rank hits |
| Total Score | Sum of scores of all aligned regions | Higher than Max Score when the hit aligns in several separate pieces |
| Acc. Len | Length of the hit sequence | Compare with your query length |
| Accession | Identifier of the hit | Click it to open the full NCBI record |
1.3.2Read the scores together
| What you see | What it suggests |
|---|---|
| Low E value, high cover, high identity | A close homolog, likely with the same function |
| Low E value, high cover, moderate identity | A more distant homolog; function may be similar but should be checked |
| Low E value, low cover | The sequences share only a region, such as a common domain |
| High identity, very short alignment | Often a chance match; check the E value |
| E value above 0.01 | Probably no meaningful similarity |
1.3.3Check the alignment
Check the alignment of your top hits in the Alignments tab (shortened example):
Query 1 MKCTNIFLTIGLVFLALLLGSQ 22
MKC+NIFL IG+VFLALLLG+Q
Sbjct 1 MKCSNIFLAIGIVFLALLLGAQ 22
| Symbol in the middle line | Meaning |
|---|---|
| A letter | Identical residue in both sequences |
+ | Different but similar residues (proteins only) |
| Blank | Different residues |
- (in Query or Sbjct) | A gap: a residue missing in one sequence |
Above each alignment, Identities, Positives and Gaps summarise these counts.
1.4Troubleshooting
- No significant hits: check that the program matches your sequence (DNA into blastn, protein into blastp), remove any organism limit, or try a broader database such as nr.
- Error about the sequence: remove numbers, spaces or non-sequence characters, and keep only one header line starting with
>. - Too many hits from the same species: use the Organism field to exclude it, or switch to refseq_protein or swissprot.
- Top hits are all “hypothetical” or “uncharacterized” proteins: search swissprot to find characterised proteins.
- The search is slow: shorter queries and smaller databases run faster; keep the page open or save the RID.
1.5References
- Camacho C. et al. (2009). BLAST+: architecture and applications. BMC Bioinformatics 10: 421.
- NCBI BLAST help and documentation
- AI in Genomics 2026, Practical 3: Compare.
2Genome annotation
Genome annotation predicts where the genes are in an assembly and assigns a function to each of them. This chapter describes three ways to annotate an assembly: Prokka on the command line (2.1), Prokka in your browser on the Galaxy EU server (2.2), and the RAST web server (2.3).
| Route | Where it runs | Set-up | Time | What you get |
|---|---|---|---|---|
| 2.1 Prokka CLI | Your computer, from a terminal | Install Prokka (conda is recommended) | 2–20 minutes for a typical genome | GFF, GenBank, FASTA and table files |
| 2.2 Prokka GUI | Galaxy EU server, in a browser | Nothing to install; a free account is optional but recommended | Depends on the server queue | The same files as the command-line version |
| 2.3 RAST | RAST server, in a browser | Account that is approved by a person | A few hours to a day | Annotation organised into subsystems, browsed in the SEED Viewer |
2.1Prokka CLI
Command-line interface
Prokka annotates a bacterial, archaeal or viral assembly in 2–20 minutes on a laptop. It takes an assembly (FASTA or FNA) as input and predicts CDS, rRNA, tRNA, tmRNA and ncRNA features, then assigns products by matching against curated protein databases.
2.1.1Prokka installation and setup
Conda is the route that works on both Linux and macOS and pulls in all 20+ dependencies (BLAST+, Prodigal, Aragorn, Barrnap, HMMER, tbl2asn). Install Miniconda first if you do not have it.
# add the channels once, in this order
conda config --add channels defaults
conda config --add channels bioconda
conda config --add channels conda-forge
# install into its own environment (recommended)
conda create -n prokka_env -c conda-forge -c bioconda prokka
conda activate prokka_env
Alternatives if you do not use conda:
# macOS, Homebrew
brew install brewsci/bio/prokka
# Ubuntu / Debian
sudo apt-get install prokka
# Docker, no local dependencies at all
docker pull staphb/prokka:latest
Then index the databases and confirm the install:
prokka --setupdb # builds the BLAST indices, run once after install
prokka --version # should print e.g. prokka 1.14.6
prokka --listdb # shows which kingdom and genus databases are available
If prokka --version fails with a Perl module error, the usual cause is a missing dependency outside conda; reinstalling inside a clean conda environment fixes it in almost every case. Remember to run conda activate prokka_env in every new terminal session.
2.1.2Annotation
Prokka takes the assembly as a positional argument: there is no input (-i) flag.
Open a terminal.
- On Ubuntu, Ctrl+Alt+T.
- On macOS, “Terminal” from “Applications › Utilities”.
Check where you are. This prints the current working directory, so you know what any relative path will be measured against.
pwdCopy the path of the folder holding your assembly. In the file manager, right-click the folder and choose “Copy path”.
- Ubuntu: Ctrl+L in Files shows the path.
- macOS Finder: right-click and hold Option, then “Copy as Pathname”.
Move into that folder. Type
cd, a space, then paste. You will see something like this:cd /home/user/projects/genome_assemblyOR drag and drop the folder onto the terminal window, which also pastes its path. Quote the path if it contains spaces.
List the files and confirm the assembly is there using
ls, which lists all the files in your directory.lsRun Prokka.
prokka contigs.fasta --outdir results_prokka --prefix my_genome --compliant
What each part does:
contigs.fasta: your assembly file.--outdir results_prokka: the folder to create for results. It must NOT already exist, or Prokka stops. If it does, add--forceto your command to overwrite.--prefix my_genome: names every output filemy_genome.gff,my_genome.faaand so on. Without it they are named after today’s date, which gets confusing fast.--compliant: enforces GenBank/ENA/DDBJ submission rules. It is shorthand for--addgenes --mincontiglen 200 --centre XXX, so it drops contigs under 200 bp and adds gene features alongside each CDS.- Optional:
--cpus 4: threads to use;--cpus 0uses every available core.
A typical bacterial genome finishes in 2–20 minutes. Progress is printed to the terminal and mirrored in the .log file.
2.1.3Optional arguments worth adding
Prokka accepts single or double dashes interchangeably (-compliant and --compliant both work). prokka --help lists everything.
| Argument | What it does | When to use it |
|---|---|---|
--force | Overwrites an existing output directory | Re-running after a failed or tweaked attempt |
--genus Escherichia | Sets the genus in the annotation metadata | Always, if you know your organism |
--species coli | Sets the species | With --genus, for cleaner headers |
--strain K12 | Sets the strain name | Multi-isolate projects |
--usegenus | Uses the genus-specific BLAST database instead of the generic one | Much better product names, if your genus is in --listdb |
--kingdom Archaea | Selects the genetic code and databases: Bacteria (default), Archaea, Viruses, Mitochondria | Anything that is not a bacterium |
--gcode 4 | Overrides the translation table | Mycoplasma, Spiroplasma and other code-4 organisms |
--gram neg | Runs SignalP for signal-peptide prediction | Secretome or surface-protein work |
--locustag ECO | Prefix for locus tags (e.g. ECO_00001) | Submission, or keeping isolates distinct |
--centre UoT | Sequencing centre ID written into the files | Replaces the placeholder XXX that --compliant inserts |
--mincontiglen 500 | Skips contigs shorter than this | Fragmented assemblies full of short junk contigs |
--rfam | Also searches Rfam for ncRNAs with Infernal | Non-coding RNA is of interest; roughly doubles runtime |
--metagenome | Tunes Prodigal for fragmented or mixed input | MAGs and metagenome bins |
--proteins ref.gbk | Annotates against your own trusted proteins first | Matching a close reference strain’s nomenclature |
--evalue 1e-09 | Tightens the similarity cutoff (default 1e-06) | Reducing spurious product assignments |
--norrna / --notrna | Skips rRNA or tRNA prediction | Speed, when only CDS matter |
--quiet | Suppresses screen output | Running inside a loop or pipeline |
A fuller command for a known organism:
prokka contigs.fasta --outdir results_prokka --prefix ecoli_K12 \
--genus Escherichia --species coli --strain K12 --usegenus \
--locustag ECO --centre UoT --compliant \
--cpus 8 --rfam
2.1.4Annotation outputs
Every file shares the --prefix you set, so with --prefix my_genome you get my_genome.gff, my_genome.faa and so on.
| File | What it contains | What you use it for |
|---|---|---|
.gff | GFF3: every predicted feature with coordinates, strand and product, plus the contig sequences at the end. | The master annotation file. Feed it to Roary, Artemis, IGV or any pangenome tool. |
.gbk | The same annotation in GenBank flat-file format, sequence and features together. | Viewing in Artemis or SnapGene, and as input to tools that expect GenBank. |
.faa | Amino-acid FASTA of every translated CDS, headers carrying the locus tag and product. | The workhorse for downstream function: BLAST, eggNOG, KEGG, InterProScan, AMR screening. |
.ffn | Nucleotide FASTA of all transcripts: CDS plus rRNA, tRNA, tmRNA and misc_RNA. | Gene-level nucleotide analyses, primer design, per-gene alignments. |
.fna | Nucleotide FASTA of the input contigs, unchanged apart from renamed headers. | The reference copy of the assembly that matches the annotation coordinates. |
.tsv | Tab-separated table: locus_tag, feature type, length, gene name, EC number, COG and product. | Open in Excel or R to search, filter and count genes. Usually the first file to look at. |
.txt | A short statistics summary: contig count, bases, and counts of CDS, rRNA, tRNA, tmRNA. | Quick sanity check: a 4 Mb genome with 300 CDS means something went wrong. |
.tbl | NCBI feature table, the plain-text intermediate used to build the .sqn. | Only touched during GenBank submission or manual feature correction. |
.sqn | ASN.1 Sequin file generated by tbl2asn, bundling sequence and annotation. | The file you actually submit to GenBank. |
.fsa | Nucleotide FASTA of the contigs with extra Sequin tags in the headers. | Input for the .sqn; ignore it unless you are submitting. |
.err | NCBI discrepancy report listing annotations that break submission rules. | Read it before submitting; it tells you exactly what a curator would reject. |
.log | Full run log: every command executed and the version of each dependency. | Reproducibility, methods sections, and debugging a run that looks odd. |
Quick checks after a run
cd results_prokka
cat my_genome.txt # feature counts
head -3 my_genome.tsv # table header and first rows
grep -c ">" my_genome.faa # number of proteins predicted
As a rough guide, a typical bacterial genome yields around 900 CDS per Mb. Far fewer usually means a fragmented assembly or the wrong --kingdom.
2.1.5Resources and references
- Seemann, T. (2014). Prokka: Rapid prokaryotic genome annotation. Bioinformatics, 30(14), 2068–2069. https://doi.org/10.1093/bioinformatics/btu153
- Prokka on GitHub: official documentation, full option list and output specification
- Prokka on Bioconda: installation recipe and current version
2.2Prokka GUI
On the Galaxy EU server
This workflow guides you through the annotation of your genome using the web-based version of Prokka, which is available on the Galaxy EU Server.
2.2.1Overview of Galaxy EU
The European Galaxy Server (usegalaxy.eu) is a public web platform for accessible, reproducible, and transparent computational research, primarily used for data analysis in fields like bioinformatics, genomics, and climate science.
Users can interact with thousands of specialized scientific tools via a graphical web interface rather than writing code.
2.2.2Running Prokka on the Galaxy EU Server
Set up your Galaxy EU account
- Open usegalaxy.eu in any browser.
- Click Login or Register in the top menu and create a free account. You can run tools anonymously, but an account keeps your histories, gives a larger storage quota and lets Galaxy email you when a job finishes.
- Start a new history for this analysis: in the “> History” panel on the right, click the + icon and rename it (click the name) to something like Prokka – isolate 12. A history is Galaxy’s project folder; every file you upload and every result appears there as a numbered dataset.
The screen has three panels: tools on the left, the working area in the centre where tool forms open, and your history on the right.
Search for Prokka
- In the search tools box at the top of the left panel, type
Prokka. - Click Prokka – Prokaryotic genome annotation. The tool form opens in the centre panel.
- Check the version button at the top of the form and use the newest one listed (1.14.6 builds at the time of writing), so results match the command-line version.
You can open the form now, but it only becomes useful once your genome is uploaded: the input dropdown lists datasets from your current history.
2.2.3Upload your genome
This is where most participants get stuck. The upload itself is easy; the trap is that Galaxy must recognise the file as fasta, or it will not appear in Prokka’s input list.
Upload the file
- Under Contigs to annotate, click on “...”. Click the Upload icon (an arrow pointing up) at the bottom left of the panel. The upload dialog opens.
- Add your assembly one of two ways:
- Choose local file: pick the contigs FASTA from your computer.
- Paste/Fetch data: paste a URL, for example an NCBI or Zenodo link, and Galaxy downloads it directly.
- Set the Type column to
fasta. It defaults to auto-detect, which usually works but sometimes labels an assembly astxt. Setting it yourself removes the guesswork. Leave Genome as?. - Click Start, then Close the dialog. You do not need to keep it open.
Galaxy decompresses .gz and .zip files automatically, so a gzipped assembly is fine as it is. Upload the assembly, not the raw reads (.fastq): Prokka annotates contigs.
Check it arrived correctly and run the tool
The dataset appears at the top of your history and changes colour as it goes: grey queued, orange running, green done, red failed.
- When it is green, this means that the upload is done, NOT the annotation. Click its name to expand it and check the number of your upload (your first upload would be 1).
- Click the selection box under Contigs to annotate and choose the number that corresponds to your upload. In case you upload multiple genomes at once, this tells the tool which genome to annotate.
- Click Run Tool at the top right of the central panel.
- Output files will start showing in your history: grey (queued), orange (running), green (complete), or red (failed).
- Once complete, you may download the output file of your choice using the save icon that is found in each green box. OR click the eye icon to preview the contents; you should see
>header lines followed by sequence.
- When the outputs turn green they are named like Prokka on data 1: gff, Prokka on data 1: faa. The file contents are identical to the command-line version, which explains what each one holds (see the output table in 2.1 Prokka CLI).
- One file at a time: click the dataset name to expand it, then the download icon (a floppy disk). Galaxy names the download after the dataset, so rename it on your computer, for example to
isolate12.gff. - Everything at once: open the history options menu (the icon at the top of the history panel) and choose Export History to File. Galaxy packages every dataset into one archive you can download. This also keeps a record of the exact tool versions and settings, which is handy for methods sections.
- Or keep working in Galaxy: the outputs are already in place for your next steps.
- You can switch on Email notification at the bottom to receive updates on your jobs via email.
2.2.4Troubleshooting
| Problem | What you see | Fix |
|---|---|---|
| Wrong format | Format shows txt or tabular; the file is missing from Prokka’s input dropdown | Click the pencil (Edit attributes) → Datatypes tab → choose fasta → Save. No re-upload needed. |
| Uploaded reads instead | Format shows fastqsanger | Prokka needs an assembly. Assemble first (for example with Shovill or SPAdes, both on Galaxy EU), or upload the contigs file instead. |
| Upload stays grey | Nothing happens for a long time | The server queue is busy. Wait a few minutes before re-uploading; duplicates only add to the queue. |
| Upload turns red | Error on the dataset | Click the bug icon to read the error. Usually a broken download link or an empty file. |
| Contig names too long | Prokka later fails with a message about contig ID length | Prokka rejects contig IDs over 37 characters, which long SPAdes names can exceed. Shorten the headers before uploading, then re-run. |
2.2.5Sources
- Galaxy EU: the server, tool search and upload
- Galaxy Training: genome annotation with Prokka: the official step-by-step tutorial, useful as a companion exercise
2.3RAST
Rapid Annotation using Subsystems Technology
The Rapid Annotation using Subsystems Technology (RAST; rast.nmpdr.org) annotates an uploaded assembly on a remote server and returns it organised into subsystems (curated functional groups of genes), which you browse in the SEED Viewer. There is nothing to install, but jobs are queued and typically take a few hours to a day.
2.3.1Create a RAST account
Accounts are approved by a human, so register a few days before you need results.
- Go to rast.nmpdr.org.
- Click Register for a RAST account in the left-hand menu.
- Fill in the form: name, institutional email, institution, and a short note on what you intend to annotate. An academic or institutional email address is approved faster than a personal one.
- Submit, then wait for the approval email. This usually arrives within a working day, occasionally longer.
- Log in from the RAST home page with the username and password you chose.
One account covers unlimited jobs, and your genomes stay private to you unless you explicitly make them public.
2.3.2Upload a new job
- Log in and click Upload New Job in the left menu.
- Choose the sequence file. Click Browse and select your assembly: a contigs FASTA (
.fasta,.fna,.fa), optionally gzipped. Upload the assembly, not raw reads; RAST does not assemble. - Click Use this data to move to the settings page.
- Fill in the organism taxonomic details.
- Taxonomy ID: enter the NCBI taxid if you know it and click the lookup button; it auto-fills the rest. Otherwise leave it and type the fields manually.
- Domain: Bacteria, Archaea or Virus.
- Genus, species, strain: as precise as you can be. If unidentified, a genus plus “sp.” and your isolate code works.
- Genetic code: 11 for most bacteria and archaea, 4 for Mycoplasma, Spiroplasma and relatives.
- Choose the annotation scheme. Select RASTtk, the current toolkit and the default. Classic RAST is kept only for reproducing old jobs.
- Set the options. The defaults are sensible; the ones worth a thought:
- Automatically fix errors: leave ticked.
- Fix frameshifts: tick it for draft assemblies, where sequencing errors create false frameshifts.
- Build metabolic model: tick it if you want a draft flux-balance model in ModelSEED.
- Backfill gaps: a second pass to catch genes missed on the first; harmless to leave on.
- Click Finish the upload. You get a job ID immediately, and an email when the job completes.
Runtime depends on the queue and on genome size. Check progress under Jobs Overview, which shows each stage with a coloured status per stage.
2.3.3Find and read the results
When the email arrives, go to Jobs Overview and click the job, or use My Jobs. Two buttons matter: View details, which opens the SEED Viewer, and Download, which exports the annotation as GenBank, GFF3, EMBL, amino-acid FASTA, nucleotide FASTA or an Excel spreadsheet.
Clicking View details lands you on the organism genome metrics: contig count, genome size, GC content, number of coding sequences, number of RNAs, and the proportion of genes placed in subsystems.
To browse the actual annotation, click Browse annotated genome in SEED Viewer.
Subsystem statistics: the two views of the same data
A subsystem is a curated set of genes that together carry out one biological process: a pathway, a complex, a transport system. The subsystem statistics panel shows them two ways, side by side.
| View | What it shows | Best for |
|---|---|---|
| Circular diagram (pie chart) | Each slice is a top-level category (amino acids, carbohydrates, virulence, and so on), sized by how many genes fall in it and colour-coded. | Seeing the shape of the genome at a glance, and spotting an unusually large category. Click a slice to filter the table to it. |
| Subsystem table | A sortable hierarchy: category → subcategory → subsystem, with the gene count in each row. | Actually finding genes. Anything you want to click through to starts here. |
The two are linked: clicking a pie slice filters the table, and the table is where you drill down. Use the pie to orient yourself, the table to work.
Subsystem coverage is typically 40–60% of genes. The rest are real genes that are simply not in a curated subsystem yet, most of them hypothetical proteins. A low percentage is not a failed annotation.
2.3.4Choose a subsystem in the table
The table is a three-level hierarchy. You expand your way down to a named subsystem, and that subsystem gives you its gene list.
- Pick a category. The top-level rows are broad: Amino Acids and Derivatives, Carbohydrates, Cell Wall and Capsule, Virulence · Disease · Defense, Membrane Transport, and so on. Click the category name, or the matching slice in the circular diagram.
- Expand to a subcategory. Inside Carbohydrates, for example, you get Central carbohydrate metabolism, Fermentation, Monosaccharides, and others.
- Click the subsystem itself. Inside Central carbohydrate metabolism you might pick Glycolysis and Gluconeogenesis or TCA Cycle. The subsystem name is the clickable link.
- Read the subsystem page. It has two parts:
- The functional roles table: each role in the pathway, and which of your genes fills it. Blank rows are informative: a missing role may mean an incomplete pathway, or a gene the annotation missed.
- The gene list, with a feature ID for each, of the form
fig|83333.1.peg.1234.pegstands for protein-encoding gene; the number is the gene’s index in your genome.
- Click a feature ID to open that single gene.
A practical shortcut for a workshop exercise: Virulence · Disease · Defense › Resistance to antibiotics and toxic compounds gives participants an immediately interesting gene list in any isolate.
2.3.5One gene: neighbourhood and sequences
Clicking a feature ID (fig|83333.1.peg.1234) opens the Annotation Overview page for that gene.
The genetic environment
Scroll to the Compare Regions panel. It draws your gene as a red arrow in the centre of its contig, with the genes flanking it on either side, and stacks the equivalent region from closely related genomes underneath.
How to read it:
- Arrow direction is the strand, so you can see at a glance which neighbours are co-oriented and might sit in one operon.
- Colour equals homology. Genes sharing a colour across rows are orthologues; your red gene’s colour group tracks it through every genome shown.
- A conserved block of the same colours in the same order across many rows is a strong hint of a functional unit: an operon or a gene cluster.
- A break in synteny (neighbours that differ from the related genomes) often marks a mobile element, a genomic island, or a horizontal transfer event.
- Hover over any arrow for its function; click it to jump to that gene’s own page. The controls above let you widen the window and change how many genomes are compared.
The FASTA sequences
On the same page, the sequence section gives both forms of the gene:
| Sequence | Where | Typical use |
|---|---|---|
| DNA (nucleotide) | The DNA sequence link or tab on the feature page | Primer design, cloning, checking the start codon and the region around it |
| Protein (amino acid) | The protein sequence link or tab, alongside it | BLAST against NCBI (see 1. Sequence alignment), domain search in InterPro or Pfam, structure prediction (see 3. Protein modelling) |
Both display as plain FASTA with the feature ID in the header, so you can copy-paste straight into BLAST. There are also links out to run BLAST directly, and to the gene’s entry in related databases.
For many genes at once, do not scrape this page: go back to the job page and use Download to export the whole annotation as amino-acid or nucleotide FASTA.
2.3.6Sources
- RAST server: registration, job submission and the SEED Viewer
- BV-BRC: the successor platform, running the same RASTtk annotation service
3Protein modelling
GetGenome supports the sequencing of your organism of interest, providing an assembly file (.fasta) and a protein sequence file (.faa) that is generated from the annotation of the genome. This workflow guides you through the steps to model your protein of interest (POI) using the .faa file and AlphaFold3 (AF3) Server.
3.1Overview
AlphaFold 3 Server is an advanced artificial intelligence model developed by Google DeepMind and Isomorphic Labs that predicts the 3D structures and interactions of life’s cellular molecules (proteins, DNA, RNA, chemical modifications, and small molecule ligands or ions).
3.2The input file
The protein sequence file (.faa) contains the protein sequences from your genome, with identifiers and, where possible, predicted functions. You do not need to identify the proteins yourself; you only need to find your protein of interest in the file. (A .faa file is one of the outputs of genome annotation; see 2. Genome annotation.)
3.2.1Open your protein sequence fileStep 1
Open the .faa file in a text editor (for example “Notepad”, “Sublime Text”). Do not use Microsoft Word or Microsoft Excel as they do not support such file formats.
3.2.2Locate your amino acid sequenceStep 2
AlphaFold 3 Server requires the primary amino acid sequence of your protein as the input. The protein file links each amino acid sequence to an identifier.
Use Find (Ctrl+F or Cmd+F) and type your protein’s name or symbol. A matching line looks like this (shortened):
contig_12 annotation gene 803242 803583 . - . ID=g04521;Name=AVR-Pik;product=AVR-Pik effector contig_12 annotation mRNA 803242 803583 . - . ID=g04521.t1;Parent=g04521Note the identifier of the mRNA line (here g04521.t1): the protein sequence is usually stored under this identifier.
If the gene has several mRNA lines (for example g04521.t1 and g04521.t2), the gene has more than one predicted version of its protein. Use the primary one: the version marked as canonical or representative in the annotation if it is marked, otherwise the first (.t1), which is usually the longest.
Copy the sequence. Copy the primary amino acid sequence, which will end before the next
>:>g04521.t1 MKCTNIFLTIGLVFLALLLGSQAEAETFLTIGLVFLALL >g03881.t1
Optional (recommended): Save the amino acid sequence as a new text file, for example my_protein.faa.
3.2.3Checkpoints
Before modelling, look at the sequence:
| Check | Good sign | Warning sign |
|---|---|---|
| First letter | M | Anything else: the annotation may have missed the start |
| Stop characters | No * or ., or only one at the very end | A * or . in the middle: the gene prediction is probably wrong or the gene is broken |
| Length | Similar to known versions of your protein in UniProt | Much shorter or longer: check the annotation or ask the GetGenome team |
3.2.4Troubleshooting
- Search finds nothing: gene names in the annotation may differ from the ones you know. Try the predicted function, or use the BLAST route:
If your protein is not named in the file, take its amino acid sequence and search it with blastp (see 1. Sequence alignment). The top hit gives you the identifier.
- Several different genes match: your protein may belong to a family with several members in the genome. Use BLAST against a known version to pick the closest one.
- Identifier not found in the proteins file: search for the gene identifier instead of the mRNA identifier (for example g04521 rather than g04521.t1); naming conventions vary.
3.3Input to AlphaFold 3 ServerStep 3
- Sign in at alphafoldserver.com with your Google account.
- In the first entity box, choose Protein as the molecule type.
- Paste your sequence without the header line (the line starting with
>). Remove any trailing*or.. - To model your protein together with other molecules, click Add entity and add each one: another protein, DNA, RNA, a ligand or an ion.
- Give the job a descriptive name.
- Leave the seed* on automatic for a first run.
- Preview the job, check that every entity and copy number is correct, then Confirm and submit job.
The server limits the total size of each job and the number of jobs per day. Both are shown on the server.
* A seed is the random starting number for a prediction: the same seed reproduces the same result, and different seeds give alternative predictions whose agreement tells you how reliable the structure is.
3.3.1Number of copies
Each entity has a Copies field. It sets how many copies of that same molecule are in the model.
Set more than one copy only when there is evidence that the protein works that way, from the literature or from known structures of related proteins. AlphaFold will build a complex from whatever copies you give it, and the result can look plausible even when that complex does not exist in nature.
3.4Understand the outputStep 4
When the job finishes, open it from the job list. The results page shows:
- the 3D model, coloured by confidence (see below);
- the pTM and, for complexes, ipTM scores;
- the green plot (see below);
- a selector to switch between the five models AlphaFold produces for each job.
Click Download to get a zip file:
| File | Content |
|---|---|
…_model_0.cif to …_model_4.cif | The five predicted structures. model_0 is the top-ranked one |
…_summary_confidences_0.json to _4.json | The overall scores of each model (pTM, ipTM, ranking score and others) |
…_full_data_0.json to _4.json | Per-residue and per-pair confidence values, including the data behind the green plot |
…_job_request.json | Exactly what you submitted. Keep it so the job can be reproduced |
The .cif files open in UCSF ChimeraX (see 4. Protein visualization).
3.5Interpreting the confidence scores
Read the confidence scores before looking at any structural detail. A model without its confidence scores should not be interpreted.
3.5.1Colours on the model (pLDDT)
The model is coloured residue by residue by how confident AlphaFold is in the local structure:
| Colour | Score | How to read it |
|---|---|---|
| Dark blue | > 90 | Very high confidence: backbone and side chains likely accurate |
| Light blue | 70–90 | Confident: the backbone is likely correct |
| Yellow | 50–70 | Low confidence: treat with caution |
| Orange | < 50 | Very low confidence: do not interpret the shape of this region |
Orange regions are often flexible or disordered parts of the protein. AlphaFold still draws them, sometimes as long loops or isolated helices, but their shape carries no information.
3.5.2pTM
pTM is a single score from 0 to 1 for the overall fold of the whole model.
| pTM | How to read it |
|---|---|
| > 0.5 | The overall fold may be correct |
| < 0.5 | The overall fold is likely wrong or uncertain |
3.5.3ipTM
ipTM appears only when the job has more than one entity or copy. It scores, from 0 to 1, how confident AlphaFold is in the way the entities are placed relative to each other.
| ipTM | How to read it |
|---|---|
| > 0.8 | The interaction between the entities is confidently predicted. |
| 0.6–0.8 | Grey zone: the interaction may be correctly or falsely predicted. Treat cautiously. |
| < 0.6 | The interaction is likely wrongly predicted. |
pTM and ipTM are less reliable for very small proteins and short chains; for these, rely more on the colours and the green plot.
3.5.4The PAE plot (Predicted Aligned Error)
The PAE plot shows how confident AlphaFold is in the relative position of every part of the model against every other part.
- Both axes list the residues of the model, in order, chain after chain. When there are several chains, lines on the plot mark where one chain ends and the next begins.
- Each point shows the expected error in the position of the residue on one axis when the model is aligned on the residue on the other axis.
- Dark green means low expected error (confident); light or white means high expected error (not confident).
How to read common patterns:
| Pattern | Meaning |
|---|---|
| Dark square along the diagonal | A confidently modelled domain |
| Two dark squares on the diagonal, light between them | Two domains each modelled confidently, but their position relative to each other is uncertain |
| Dark blocks off the diagonal, between two chains | The two chains are confidently placed relative to each other; supports the predicted interface |
| Light blocks between two chains | AlphaFold is not confident how the chains are arranged, whatever the model looks like |
The colours on the model and the green plot answer different questions. A region can be dark blue on the model, meaning its local structure is confident, while the green plot shows its position relative to another domain or chain is uncertain.
3.5.5Reading the scores together
| What you see | What it suggests |
|---|---|
| Mostly blue model, pTM > 0.5 | The fold is likely reliable |
| Blue chains, high ipTM, dark green between chains | The complex is confidently predicted |
| Blue chains, low ipTM, light green between chains | Each chain is folded confidently, but the interaction is not supported |
| Mostly orange model | Possible disordered protein, missing partner, or wrong sequence (go back to Step 2) |
| The five models look very different | The prediction is uncertain, even if one model scores well |
A confident model is a well-founded hypothesis about your protein’s structure, not proof. Use it to decide which experiments are worth doing.
3.6Troubleshooting
- “Invalid sequence”: a header line, a
*or., spaces, or letters that are not amino acids were pasted. - Job too large: model a single domain, or reduce the number of copies or partners.
- Daily limit reached: the counter resets; submit the job the next day.
- Everything is orange: check the sequence (Step 2 checkpoints) before re-running.
3.7References
- Abramson J. et al. (2024). Accurate structure prediction of biomolecular interactions with AlphaFold 3. Nature.
- Maqbool A. et al. (2015). Structural basis of pathogen recognition by an integrated HMA domain in a plant NLR immune receptor. PMID 26304198.
- Posbeyikian A., Toghani A., McClelland, Kamoun S., Sugihara Y., Contreras (2026). Fold on Tight: Ten Things to Know About AI Structure Prediction in Plant Immunity.
4Protein visualization
UCSF ChimeraX is a free program to visualise and analyse 3D protein structures. This workflow guides you through installing ChimeraX and learning its basic commands.
4.1InstallationStep 1
Download ChimeraX from the official download page, choosing the version for your operating system. It is free for academic and non-commercial use: accept the licence to start the download. Install the latest version, as AlphaFold 3 files need a recent one.
| Operating system | How to install |
|---|---|
| Windows | Run the .exe installer and follow the prompts. |
| macOS | Open the .dmg file and drag ChimeraX into Applications. Apple silicon (M1 and later) and Intel Macs have separate downloads. |
| Linux | Install the .deb (Ubuntu, Debian) or .rpm (Fedora, Red Hat) package. |
Open ChimeraX and type version in the command line at the bottom of the window: the version number appears in the Log.
4.1.1Troubleshooting
- Installation is blocked: institutional laptops often need administrator rights; ask your IT department.
- macOS says the app cannot be opened: go to System Settings → Privacy & Security and click Open Anyway.
4.2Using ChimeraXStep 2
ChimeraX is used through menus and through commands typed in the command line. Every command and its result appear in the Log.
| Part of the window | What it does |
|---|---|
| Command line (bottom) | Type a command and press Enter. |
| Log | Shows commands, results and error messages. |
| Toolbar tabs (Home, Molecule Display…) | One-click actions, such as colouring by chain or taking a picture. |
| Models panel | Lists open models; show, hide or close them. |
| Sequence icon | Shows the sequence of each chain, linked to the 3D structure. |
4.2.1Open a structure
Open a structure from the Protein Data Bank (PDB) with its four-character ID:
open 7B1I
Click [more info...] in the Log to see its metadata, such as the paper title and the experimental method.
4.2.2Point to parts of a structure
Use these symbols, combined from left to right (for example #1/b:1-10):
| Symbol | Level | Example |
|---|---|---|
# | Model | #1 = the first model opened |
/ | Chain | /A = chain A |
: | Residue | :51 = residue 51; :1-10 = residues 1 to 10; :1,5,20 = residues 1, 5 and 20 |
@ | Atom | @ca = alpha carbon atoms |
4.2.3Select, colour and display
| Command | What it does |
|---|---|
sel #1/c | Selects chain C |
color sel gray | Colours the selection gray |
color #1/c #2B89AD | Colours chain C with a hex colour code |
color #1/b:3-9 blue | Colours residues 3 to 9 of chain B |
color #1 bychain | Gives each chain its own colour |
surface #1/c | Shows the surface of chain C |
hide #1 surfaces | Hides all surfaces |
show #1/a:51 atoms | Shows the atoms (side chain) of residue 51 of chain A |
view | Centres the structure in the window |
help color | Opens the documentation for any command |
Close the structure before the next step:
close #1
4.3Open your AlphaFold 3 modelStep 3
Unzip your AlphaFold 3 download (see 3. Protein modelling) and drag …_model_0.cif (the top-ranked model) into the ChimeraX window, or type its path:
open ~/Downloads/fold_my_job/fold_my_job_model_0.cif
4.3.1Colour by pLDDT
Blue = confident, orange = low confidence:
color bfactor #1 palette alphafold
4.3.2Load the PAE plot
Go to Tools → Structure Prediction → AlphaFold Error Plot and select your model. ChimeraX usually finds …_full_data_0.json; if not, click Browse, select it, then Open. The same, by command:
alphafold pae #1 file ~/Downloads/fold_my_job/fold_my_job_full_data_0.json
The PAE plot is interactive: drag a box over a region of the plot and it is highlighted on the 3D model. The plot window also recolours the model by pLDDT or by PAE domains.
4.4Analyse the interfaceStep 4
4.4.1Show contacts
Show contacts between chain A and chain B closer than 3.5 Å. Each contact is drawn as a line (pseudobond) coloured by PAE, from blue (confident) to red (not confident). The PAE plot must be loaded first (Step 3).
alphafold contacts #1/a to #1/b distance 3.5
To look at one residue only, here residue 20 of chain B within 4.5 Å of chain A:
alphafold contacts #1/b:20 to #1/a distance 4.5
4.4.2Export the contacts
Export the contacts to a text file, to sort them in Excel:
alphafold contacts #1/a to #1/b distance 3.5 outputFile AF3-contacts.txt
Each row is a pair of residues in contact and its PAE, for example /A:81 /B:54 17.50.
| PAE (Å) | Confidence |
|---|---|
| 0–10 | Confident |
| 10–15 | Grey zone |
| > 15 | Poor |
Residues in confident contacts are candidates to mutate when testing the interaction in the lab.
4.5Save your workStep 5
ChimeraX does not save automatically.
| Command | What it saves |
|---|---|
save ~/Desktop/my_session.cxs | The whole session; reopen it with open |
save ~/Desktop/image.png supersample 3 | A high-quality picture of the current view |
To record a 360° spin movie, run these four commands in order:
movie record
turn y 2 180
wait 180
movie encode ~/Desktop/spin.mp4
Pictures and movies can also be taken from Home → Images. Type pwd to see the folder where files are saved. On Windows, you can also write full paths, such as C:\Users\yourname\Desktop\image.png.
4.6Troubleshooting
- Command not recognised: check the spelling, or type
helpfollowed by the command name. - Wrong residues coloured: check the model number (#), chain ID (/) and residue numbers in the Models panel and the Sequence viewer.
- The PAE plot does not load: select the full_data file with the same number as the model (model_0 with full_data_0).
- The contacts command gives an error: load the PAE plot first (Step 3).
- File not found: the path is wrong; drag the file into ChimeraX instead.
4.7References
- Meng E.C. et al. (2023). UCSF ChimeraX: Tools for structure building and analysis. Protein Science 32: e4792.
- ChimeraX User Guide
- AI in Genomics 2026, Practical 2: Interpret (slides).
5Structure alignment
Foldseek compares a protein structure with millions of known and predicted structures and finds the most similar ones, even when their sequences are unrelated. This workflow guides you through preparing your structure, running a search, reading the results, and aligning several structures with FoldMason.
5.1Before you startStep 1
Get everything you can from the sequence first: run BLAST (see 1. Sequence alignment) and InterProScan. Use Foldseek when the sequence gives few or no answers.
Foldseek needs a structure file in PDB or CIF format, either an experimental structure or a prediction (for example, …_model_0.cif from AlphaFold 3; see 3. Protein modelling). Your results are only as good as the structure you put in, so check it first:
| Check | Why | How |
|---|---|---|
| Confidence | Low-confidence regions give misleading matches | Colour by pLDDT in ChimeraX (see 4. Protein visualization); remove regions below 50 if they are long |
| One chain | A complex mixes the signals of several proteins | Keep only your protein’s chain |
| Format | Foldseek reads PDB and mmCIF files | Convert in ChimeraX if needed |
To keep one chain and save it as a PDB file in ChimeraX (here, removing chain B):
open model_0.cif
delete /B
save my_protein.pdb
5.2Searching with FoldseekStep 2
Go to Foldseek Search.
- Upload your structure with UPLOAD PDB, or use LOAD ACCESSION to load a structure by its identifier.
Choose the databases. Keep all of them for the broadest search, or untick the ones you do not need:
Database Contains PDB100 Experimental structures from the Protein Data Bank AlphaFold/Swiss-Prot Predicted structures of manually reviewed proteins AlphaFold/UniProt50 Predicted structures covering most known proteins AlphaFold/Proteome Predicted structures from selected model organisms CATH50 Classified protein domains Other databases Predicted structures from metagenomes, viruses and other sources Choose the mode:
Mode Compares Use it to 3Di/AA Local similarity, using structure and sequence together Run a fast, sensitive search for related proteins (default) TM-align Overall (global) structural similarity Test whether hits share the whole fold LoL-align Local structural similarity Find shared regions that suggest a common origin - Optional settings. Use Taxonomic filter to restrict hits to a group of organisms. Tick Iterative search to find more distant relatives.
- Click SEARCH. Results usually appear within a few minutes. Bookmark the page, or download the results to reload them later with UPLOAD PREVIOUS RESULTS.
5.3Reading the Foldseek resultsStep 3
Results are listed per database, best hits first. Click a hit to see its alignment with your query and the two structures superposed.
| Column | What it means | How to read it |
|---|---|---|
| TM-score | Overall structural similarity, from 0 to 1 (TM-align mode) | Below 0.2: unrelated; above 0.5: generally the same fold |
| Prob. | Probability that the hit is a true homolog | Close to 1: likely related |
| E-Value | Number of hits this good expected by chance | The smaller, the more significant |
| Seq. Id. | Percentage of identical residues in the alignment | Low identity with a high TM-score suggests a hidden relationship |
| Pos. in Query / Target | Where the alignment starts and ends in each protein | Shows whether the whole protein or only part of it matches |
5.3.1Read the scores together
| What you see | What it suggests |
|---|---|
| High TM-score, high sequence identity | A close homolog; BLAST should also find it |
| High TM-score, low sequence identity | Same fold without detectable sequence similarity |
| High score over only part of the query | The proteins share a domain or region |
| TM-score between 0.2 and 0.5 | Uncertain; check the superposition and other tools |
A similar structure suggests, but does not prove, a shared origin or function. No single tool gives the whole answer: combine Foldseek with BLAST, domain annotation and the literature.
5.4Aligning many structures with FoldMasonStep 4
To compare several structures at once, for example your query and its best hits, use FoldMason, on the same website.
- Upload at least two structures (PDB or mmCIF): click CLICK TO SELECT FILES or drag and drop them.
- Click ALIGN.
The result is a structure-based multiple alignment: residues that occupy the same position in 3D are aligned, even when the sequences differ. Look for columns that are conserved across all structures, such as cysteines or glycines, which often hold the fold together.
5.5Troubleshooting
- The upload is rejected: check the file is PDB or mmCIF; re-save it from ChimeraX.
- Hits match a partner protein instead of yours: the file contains several chains; keep only your protein (Step 1).
- Many weak hits: switch to TM-align mode and focus on hits with TM-score above 0.5.
- No hits: try 3Di/AA mode with iterative search, and keep all databases selected.
- Hits only match a disordered or low-confidence region: remove low-pLDDT regions and search again.
5.6References
- van Kempen M. et al. (2024). Fast and accurate protein structure search with Foldseek. Nature Biotechnology 42: 243–246.
- Xu J. and Zhang Y. (2010). How significant is a protein structure similarity with TM-score = 0.5? Bioinformatics 26: 889–895.
- Foldseek on GitHub
- AI in Genomics 2026, Practical 3: Compare (slides).
6Glossary
Key terms used across the AI in Genomics workflows, grouped by topic and listed from A to Z within each topic.
Filter by topic
Genomics fundamentals13 terms
- Accessory genome
- Genes present in only some individuals or strains of a species; often linked to adaptation, such as virulence or resistance.
- Chromosome
- A long DNA molecule carrying many genes, packaged with proteins; each species has a typical number of chromosomes.
- Contig
- A continuous stretch of sequence assembled from overlapping reads, with no gaps.
- Core genome
- Genes shared by all individuals or strains of a species; they usually carry essential functions.
- GC content
- The percentage of G and C bases in a DNA sequence; it varies between species and between genome regions.
- Genome
- The complete DNA of an organism, including all its genes and non-coding sequences.
- Genome size
- The total length of an organism's genome, usually given in base pairs (bp, Mb or Gb).
- Haplotype
- A set of variants inherited together on the same copy of a chromosome.
- Pangenome
- The full set of genes found across all individuals of a species: the core genome plus the accessory genome.
- Ploidy
- The number of complete chromosome sets in a cell: haploid (one), diploid (two) or polyploid (more than two).
- Reference genome
- A representative genome assembly of a species, used as a standard to map reads and compare other genomes.
- Scaffold
- Contigs ordered and oriented relative to each other, with gaps of estimated size between them.
- Synteny
- Conservation of gene order along chromosomes between species or genomes.
Genome assembly & quality9 terms
- Completeness
- How much of the expected gene content is present in an assembly, often estimated with BUSCO.
- Contamination
- Sequences in an assembly that come from another organism, such as bacteria or the host.
- Coverage
- The fraction of the genome covered by at least one read; also often used to mean average depth.
- De novo assembly
- Building a genome sequence from reads without using a reference genome.
- Depth
- The average number of reads covering each position of the genome, for example 30×.
- k-mer
- A subsequence of length k; k-mer counts are used to assemble genomes and estimate genome size.
- L50
- The smallest number of contigs that together contain half of the assembly length; lower is better.
- N50
- The length of the shortest contig among the longest contigs that together make up half of the assembly; higher is better.
- Reads
- DNA sequences produced by a sequencing machine, short or long, which are assembled or mapped to a reference.
Genome annotation11 terms
- CDS (coding sequence)
- The part of a gene that is translated into protein, from the start codon to the stop codon.
- Exon
- A part of a gene kept in the mature mRNA after splicing; it can include coding and untranslated sequence.
- Functional annotation
- Assigning predicted functions to genes, such as names, domains or GO terms, usually by similarity to known proteins.
- Gene model
- The predicted structure of a gene: its position, exons, introns, CDS and UTRs.
- Gene prediction
- Identifying the position and structure of genes in a genome with software, experimental evidence, or both.
- Intron
- A non-coding part of a gene, removed from the mRNA by splicing.
- ORF (open reading frame)
- A stretch of DNA from a start codon to a stop codon, with no stop codon in between, that may encode a protein.
- Pseudogene
- A gene copy that has lost its function, often through mutations such as premature stop codons.
- Structural annotation
- Identifying where genes and their parts (exons, introns, CDS, UTRs) are located in a genome.
- Transcript/protein prediction
- Deriving the mRNA and protein sequences of each gene from its gene model.
- UTR (untranslated region)
- The parts of an mRNA before (5′ UTR) and after (3′ UTR) the coding sequence, which are not translated.
Sequence analysis13 terms
- Alignment
- Arranging two or more sequences so that corresponding residues line up, with gaps where needed.
- BLAST
- A tool that searches databases for sequences similar to a query and scores how significant each match is.
- Conservation
- How unchanged a position or region remains across related sequences; conserved regions are often important for function.
- Domain
- A region of a protein that folds and functions largely on its own and is found in different proteins.
- Global alignment
- An alignment over the full length of both sequences.
- HMM (hidden Markov model)
- A statistical model that describes a sequence as a series of hidden states; widely used to detect patterns in sequences.
- Local alignment
- An alignment of only the most similar regions of two sequences; the approach used by BLAST.
- Motif
- A short, recurring sequence pattern, often with a specific function, such as a binding site.
- Multiple sequence alignment (MSA)
- An alignment of three or more sequences, used to find conserved positions and build phylogenetic trees.
- Profile HMM
- An HMM built from a multiple sequence alignment, describing which residues each position of a family allows; used by Pfam.
- Sequence identity
- The percentage of aligned positions that have identical residues.
- Sequence similarity
- The percentage of aligned positions that have identical or chemically similar residues.
- Whole-genome alignment
- Aligning entire genomes to find conserved regions, rearrangements and variants.
Protein structure & structural bioinformatics6 terms
- ipTM
- AlphaFold score (0 to 1) of confidence in the interfaces between entities; above 0.8 indicates a confident complex.
- PAE (predicted aligned error)
- AlphaFold estimate, in Å, of the error in the relative position of two residues; low values mean confident placement.
- pLDDT
- AlphaFold per-residue confidence score (0 to 100) in the local structure; above 70 is generally reliable.
- pTM
- AlphaFold score (0 to 1) of confidence in the overall fold; above 0.5 suggests the fold may be correct.
- RMSD (root mean square deviation)
- The average distance, in Å, between matching atoms of two superposed structures; lower means more similar.
- TM-score
- A score (0 to 1) of the overall similarity of two structures; above 0.5 generally indicates the same fold.
Population genomics9 terms
- Admixture
- Mixing of genetic ancestry from previously separated populations; also the name of a tool that estimates it.
- FST
- A measure (0 to 1) of genetic differentiation between populations: 0 means none, 1 means completely separated.
- Genetic diversity
- The amount of genetic variation within a population or species.
- GWAS (genome-wide association study)
- A study that links genetic variants across the genome to a trait or disease.
- Linkage disequilibrium (LD)
- The non-random association of alleles at different positions, often because they are close on a chromosome.
- Nucleotide diversity (π)
- The average number of differences per site between two sequences taken from a population.
- PCA (principal component analysis)
- A method that summarises many variables into a few axes; used to visualise population structure.
- Population structure
- Genetic differences between groups within a species, caused by limited mixing between them.
- Selective sweep
- Reduced diversity around a beneficial variant that rose quickly in frequency under selection.
Genomic data formats16 terms
- BAM
- The compressed, binary version of SAM; smaller and faster to read.
- CRAM
- A highly compressed alignment format that stores reads relative to a reference genome.
- CSV (comma-separated values)
- A table stored as plain text, with columns separated by commas.
- FASTA
- A text format for sequences: a header line starting with > followed by the DNA or protein sequence.
- FASTQ
- A text format for sequencing reads, storing each read's sequence and the quality score of each base.
- GFF (General Feature Format)
- A tab-separated format describing genome features, such as genes and exons, and their positions.
- GTF (Gene Transfer Format)
- A format similar to GFF, commonly used for gene and transcript annotations.
- JSON
- A text format that stores structured data as key–value pairs; AlphaFold confidence scores are saved as JSON.
- Metadata
- Information that describes data, such as sample origin, sequencing method or date.
- mmCIF
- The current standard format for 3D structures of molecules, replacing PDB format; AlphaFold 3 models are .cif files.
- Newick
- A text format that represents phylogenetic trees with nested parentheses, branch lengths and labels.
- PDB format
- An older text format for the 3D atomic coordinates of molecular structures.
- SAM (Sequence Alignment/Map)
- A text format that stores reads aligned to a reference genome.
- TSV (tab-separated values)
- A table stored as plain text, with columns separated by tabs.
- VCF (Variant Call Format)
- A text format listing genetic variants, such as SNPs and indels, and the genotype of each sample.
- YAML
- A human-readable text format for structured data, often used for configuration files.
Command line / HPC16 terms
- Bash
- The most common shell on Linux and macOS, and a scripting language to automate commands.
- CPU
- The processor that runs most computations; cluster jobs request a number of CPU cores.
- Environment
- An isolated set of software and versions, such as a conda environment, used to run analyses reproducibly.
- GPU
- A processor specialised in parallel computation; needed to run AI models such as AlphaFold efficiently.
- Job
- A task submitted to a computing cluster, which runs when the requested resources are available.
- Linux
- A free, open-source, Unix-like operating system used on most servers and computing clusters.
- Modules
- A system on computing clusters to load and unload pre-installed software, for example module load blast.
- RAM
- Working memory used during computation; large genomes and AI models need more RAM.
- Resources
- The CPUs, GPUs, memory and time a job requests from a computing cluster.
- Scratch
- Fast, temporary storage on a cluster for running jobs; files are often deleted automatically after a set time.
- Shell
- A program that reads commands typed in the terminal and runs them.
- Slurm
- A job scheduler that queues and manages jobs on computing clusters.
- SSH (Secure Shell)
- A secure way to log in to a remote computer, such as a cluster, from your terminal.
- Storage
- Long-term disk space for data and results, usually backed up, unlike scratch.
- Terminal
- A window where you type text commands to control a computer.
- Unix
- A family of operating systems on which Linux and macOS are based.
Programming & data analysis12 terms
- API (application programming interface)
- A defined way for programs to request data or services from another program or database, such as NCBI.
- Biopython
- A Python library for biological computation, such as reading sequence files and querying databases.
- Data frame
- A table of data with named columns, each holding one type of value; central to pandas and R.
- Git
- A version control system that records changes to files and lets you return to earlier versions.
- GitHub
- A website that hosts Git repositories, for sharing and collaborating on code.
- Jupyter
- Notebooks that combine code, results, plots and notes in one document, run in a web browser.
- pandas
- A Python library for working with tables (data frames).
- Plotting
- Showing data as graphs or charts, for example with matplotlib in Python or ggplot2 in R.
- Python
- A general-purpose programming language widely used in bioinformatics and machine learning.
- R
- A programming language for statistics and data visualisation, widely used in biology.
- Regular expression
- A pattern used to search or match text, for example to find all sequence identifiers in a file.
- tidyverse
- A collection of R packages for cleaning, transforming and plotting data, including dplyr and ggplot2.
AI/ML in genomics5 terms
- Deep learning
- Machine learning that uses neural networks with many layers; the basis of tools such as AlphaFold.
- Generative adversarial network (GAN)
- Two neural networks trained against each other: one generates data, the other judges whether it is real.
- Machine learning
- Methods that let computers learn patterns from data to make predictions, without being explicitly programmed.
- Natural language processing (NLP)
- AI methods for understanding and generating text; similar methods are used to model DNA and protein sequences.
- Reinforcement learning
- Machine learning in which a model learns by trial and error, guided by rewards.
Databases & resources9 terms
- ENA (European Nucleotide Archive)
- The European database of nucleotide sequences and sequencing reads, run by EMBL-EBI.
- GenBank
- NCBI's public database of nucleotide sequences submitted by researchers.
- InterPro
- A database that classifies proteins into families and predicts their domains by combining several databases, including Pfam.
- NCBI
- The US National Center for Biotechnology Information, which hosts major databases and tools such as GenBank and BLAST.
- PDB (Protein Data Bank)
- The archive of experimentally determined 3D structures of proteins and nucleic acids.
- Pfam
- A database of protein families and domains, each described by a profile HMM; now part of InterPro.
- RefSeq
- NCBI's curated, non-redundant set of reference sequences for genomes, transcripts and proteins.
- SRA (Sequence Read Archive)
- NCBI's database of raw sequencing reads.
- UniProt
- The main database of protein sequences and functions; its Swiss-Prot section is manually reviewed.
No terms match your search. Try another word or choose All topics.