← All tools
Structure prediction

AlphaFold 3

(Work in Progress)

Structure prediction for proteins, DNA/RNA, ligands, ions, and modified residues. Supports co-folding.

Google DeepMind and Isomorphic Labs' AlphaFold 3 predicts the joint all-atom structure of biomolecular complexes: proteins, DNA, RNA, ligands (by CCD code, SMILES, or a user-provided CCD), ions, post-translational and nucleotide modifications, glycans, and covalent bonds. Here, MSAs come from the ColabFold MMseqs2 server by default, so the ~630 GB genetic databases are not needed; its own local data pipeline is available for an operator who has installed them.

No jobs before this; it will run immediately.

Ready
Choose an example to fill in its settings and input structures.

Input

The input's name; a sanitised version of it names the output directory and files.
Add one box per unique entity. Chain IDs are uppercase letters (A, or A, B for two copies); left blank, unused IDs are assigned. A ligand is a SMILES string, or one or more CCD codes (CCD_ATP; CCD_NAG, CCD_NAG, CCD_BMA for a multi-component ligand such as a glycan, joined with Bonded atom pairs). SMILES ligands cannot take part in bonds. An ion is its CCD code (MG, ZN, CA). Modifications take a CCD code at a 1-based position. MSA paths are optional precomputed A3M files on the runner; a protein given only one of the two runs with the other empty. Cyclic chains are not supported: AlphaFold 3 does not accept bonds within or between polymer chains.
Comma-separated integer random seeds (modelSeeds); at least one is required. The model runs once per seed, and each run produces Diffusion samples structures.
Native bondedAtomPairs: covalent bonds as pairs of [entity ID, 1-based residue index, atom name], e.g. [[["A", 12, "SG"], ["B", 1, "C25"]]] for a covalent ligand, or the bonds between the components of a multi-CCD ligand. Atom names come from the CCD; bonds within or between polymer chains are not supported.
Native userCCD: chemical components in CCD mmCIF format, defining custom ligands or modified residues (referenced by component ID from a ligand's CCD codes or a modification), with named atoms for bonds and ideal coordinates for when RDKit cannot generate a conformer. Avoid underscores in custom component IDs.
A complete AlphaFold 3 input: one object in the alphafold3 dialect (name, modelSeeds, sequences of protein, rna, dna and ligand entities, optional bondedAtomPairs and userCCD, dialect "alphafold3", version 1-4), or an AlphaFold Server job list, which AlphaFold 3 converts. In this form: Bio Web writes this text to a file on the runner, so any unpairedMsaPath, pairedMsaPath, mmcifPath, or userCCDPath inside it must be an absolute path available there. Entities that already carry MSAs or templates keep them.
An AlphaFold 3 input JSON file, in the alphafold3 dialect or as an AlphaFold Server job list. In this form: Choose a file. Its contents are read here and sent as the input document.

MSAs and templates

Where MSAs and templates come from for entities whose input does not already give them. ColabFold: each distinct protein sequence is sent to the MSA server below; in a complex of several distinct proteins the server's paired alignment is placed first in each chain's unpaired MSA, as the input docs recommend for custom pairing. RNA runs MSA-free and no templates are used. None: every chain runs from its sequence alone (private sequences, offline runs), at a large cost in accuracy. AlphaFold 3 data pipeline: Jackhmmer, Nhmmer and Hmmsearch over the ~630 GB genetic databases, which must be installed on the runner.
The ColabFold MMseqs2 API the ColabFold option sends protein sequences to.
YYYY-MM-DD (--max_template_date). The data pipeline ignores templates released after it, and only CCD components released before it may fall back to their model coordinates when RDKit cannot generate a conformer. The default is the AlphaFold 3 paper's date.

Inference

Structures sampled per seed (--num_diffusion_samples). The default is 5.
Trunk recycling iterations (--num_recycles). The default is 10.
--num_seeds: run this many seeds, counting up from the input's single seed, which must then be the only one it gives. Blank uses the input's seeds as they are.
--conformer_max_iterations: raise RDKit's conformer search limit when a ligand fails with "Failed to construct RDKit reference structure". Blank keeps RDKit's default.

Output

Runtime

--flash_attention_implementation. Choose XLA for a compute capability 7.x GPU (V100, T4, RTX 20xx); the runner then also sets the XLA flag AlphaFold 3 requires for those devices.
--jax_backend. CPU inference needs the XLA attention implementation.
Zero-based JAX device index (--gpu_device), after any CUDA_VISIBLE_DEVICES filtering.
--buckets: strictly increasing, comma-separated token counts that inputs are padded to, so a compiled model can be reused. Blank keeps AlphaFold 3's defaults (256 to 5120).
Ready