← API reference
Structure prediction

AlphaFold 3 API

Structure prediction for proteins, DNA/RNA, ligands, ions, and modified residues. Supports co-folding.

Google DeepMind and Isomorphic Labs' AlphaFold 3 predicts the joint all-atom structure of biomolecular complexes: proteins, DNA, RNA, ligands (by CCD code, SMILES, or a user-provided CCD), ions, post-translational and nucleotide modifications, glycans, and covalent bonds. Here, MSAs come from the ColabFold MMseqs2 server by default, so the ~630 GB genetic databases are not needed; its own local data pipeline is available for an operator who has installed them.

Input method

Set input_mode to choose how inputs are supplied. Only fields for the selected method are used.

Example presets

Send {"preset": "quick_start_2pv7"} to load an example and its bundled inputs. Add other request fields to override its settings. The schema includes all presets under tool.presets.

presetDescription
quick_start_2pv7The README's first prediction, fold_input.json: 2PV7, one 298-residue protein entity as a homodimer (chains A and B), with one seed. Source ↗
ubiquitin_monomerThe smallest official example: human ubiquitin, one 76-residue protein chain. Source ↗
barnase_barstarA heterodimer of barnase (110 aa) and its inhibitor barstar (89 aa). With the ColabFold MSA option the two chains get a paired alignment, placed first in each chain's MSA. Source ↗
tetr_homodimerThe tetracycline repressor (218 aa) as a homodimer: one protein entity with two chain IDs. Source ↗
tetr_dimer_tetracyclineThe TetR homodimer with two copies of tetracycline, given by CCD code (TAC). Source ↗
tetr_dimer_dnaThe TetR homodimer bound to its 20 bp tetO2 operator, as two complementary DNA strands. Source ↗
calmodulin_4calciumHuman calmodulin with four calcium ions, one per EF hand. AlphaFold 3 treats ions as ligands (CCD code CA), here one entity with four IDs. Source ↗
streptavidin_biotin_smilesCore streptavidin with D-biotin given as a SMILES string rather than a CCD code. Source ↗
kras_g12c_sotorasibKRAS4B G12C with sotorasib (CCD MOV), covalently bonded to Cys12 through bondedAtomPairs. Source ↗
rnaseb_glycosylatedBovine RNase B with a Man3GlcNAc2 glycan at Asn34: a five-component CCD ligand whose linkages and attachment are given as bondedAtomPairs. Source ↗
erk2_phosphorylatedHuman ERK2 (360 aa) with activation-loop phosphothreonine (TPO) and phosphotyrosine (PTR) as protein modifications. Source ↗
u1a_rna_hairpinThe U1A RRM1 domain with the U1 snRNA stem-loop II hairpin. The RNA runs MSA-free unless the local data pipeline is used. Source ↗
modified_rnaA 25-nt tRNA-like fragment with pseudouridine, 5-methylcytidine, and 2'-O-methylguanosine modifications. RNA only, so no MSA server is contacted. Source ↗
methylated_dnaA 20 bp DNA duplex with 5-methylcytosine at each CpG, on both strands. DNA only, so no MSA server is contacted. Source ↗
Request body

Fields

The same names the web form posts. See the field type table for what each kind means over HTTP.

Name Type Required Default Description
job_name string text no af3-job Job name The input's name; a sanitised version of it names the output directory and files. up to 80 characters.
sequence_molecules string (JSON array) molecule_builder yes [{"type": "protein", "id": "A", "count": 1, "sequence": "MQIFVKTLTGKTITLEVEPSDTIENVKAKIQDK … Molecules Add one box per unique entity. Chain IDs are uppercase letters (A, or A, B for two copies); left blank, unused IDs are assigned. A ligand is a SMILES string, or one or more CCD codes (CCD_ATP; CCD_NAG, CCD_NAG, CCD_BMA for a multi-component ligand such as a glycan, joined with Bonded atom pairs). SMILES ligands cannot take part in bonds. An ion is its CCD code (MG, ZN, CA). Modifications take a CCD code at a 1-based position. MSA paths are optional precomputed A3M files on the runner; a protein given only one of the two runs with the other empty. Cyclic chains are not supported: AlphaFold 3 does not accept bonds within or between polymer chains. One of: protein, ligand, dna, rna, ion. A JSON array, sent as a string. See molecule entries.
model_seeds string text yes 1 Model seeds Comma-separated integer random seeds (modelSeeds); at least one is required. The model runs once per seed, and each run produces Diffusion samples structures. up to 400 characters.
bonded_atom_pairs string textarea no — Bonded atom pairs (JSON, optional) Native bondedAtomPairs: covalent bonds as pairs of [entity ID, 1-based residue index, atom name], e.g. [[["A", 12, "SG"], ["B", 1, "C25"]]] for a covalent ligand, or the bonds between the components of a multi-CCD ligand. Atom names come from the CCD; bonds within or between polymer chains are not supported. up to 200000 characters.
user_ccd string textarea no — User-provided CCD (mmCIF, optional) Native userCCD: chemical components in CCD mmCIF format, defining custom ligands or modified residues (referenced by component ID from a ligand's CCD codes or a modification), with named atoms for bonds and ideal coordinates for when RDKit cannot generate a conformer. Avoid underscores in custom component IDs. up to 2000000 characters.
input_json string textarea yes { "name": "ubiquitin_monomer", "sequences": [ { "protein": { "id": " … AlphaFold 3 input (JSON) A complete AlphaFold 3 input: one object in the alphafold3 dialect (name, modelSeeds, sequences of protein, rna, dna and ligand entities, optional bondedAtomPairs and userCCD, dialect "alphafold3", version 1-4), or an AlphaFold Server job list, which AlphaFold 3 converts. In this form: Bio Web writes this text to a file on the runner, so any unpairedMsaPath, pairedMsaPath, mmcifPath, or userCCDPath inside it must be an absolute path available there. Entities that already carry MSAs or templates keep them. up to 5000000 characters.
inputs_file string file no — Input file (JSON) An AlphaFold 3 input JSON file, in the alphafold3 dialect or as an AlphaFold Server job list. In this form: Choose a file. Its contents are read here and sent as the input document. file types .json.
msa_source string select no colabfold MSAs and templates Where MSAs and templates come from for entities whose input does not already give them. ColabFold: each distinct protein sequence is sent to the MSA server below; in a complex of several distinct proteins the server's paired alignment is placed first in each chain's unpaired MSA, as the input docs recommend for custom pairing. RNA runs MSA-free and no templates are used. None: every chain runs from its sequence alone (private sequences, offline runs), at a large cost in accuracy. AlphaFold 3 data pipeline: Jackhmmer, Nhmmer and Hmmsearch over the ~630 GB genetic databases, which must be installed on the runner. One of: colabfold, none, databases.
msa_server_url string text no https://api.colabfold.com MSA server The ColabFold MMseqs2 API the ColabFold option sends protein sequences to. up to 400 characters.
max_template_date string text no 2021-09-30 Maximum template date YYYY-MM-DD (--max_template_date). The data pipeline ignores templates released after it, and only CCD components released before it may fall back to their model coordinates when RDKit cannot generate a conformer. The default is the AlphaFold 3 paper's date. up to 10 characters.
resolve_msa_overlaps boolean checkbox no true Deduplicate unpaired against paired MSAs --resolve_msa_overlaps, the paper's behaviour. Turned off automatically when the ColabFold option builds paired alignments, since deduplication would break their row pairing.
num_diffusion_samples number number no 5 Diffusion samples Structures sampled per seed (--num_diffusion_samples). The default is 5. minimum 1, maximum 64, step 1.
num_recycles number number no 10 Recycles Trunk recycling iterations (--num_recycles). The default is 10. minimum 1, maximum 100, step 1.
num_seeds number number no — Number of seeds (optional) --num_seeds: run this many seeds, counting up from the input's single seed, which must then be the only one it gives. Blank uses the input's seeds as they are. minimum 2, maximum 1000, step 1.
conformer_max_iterations number number no — RDKit conformer iterations (optional) --conformer_max_iterations: raise RDKit's conformer search limit when a ligand fails with "Failed to construct RDKit reference structure". Blank keeps RDKit's default. minimum 1, maximum 100000, step 1.
fix_standalone_glycans boolean checkbox no false Keep leaving atoms of standalone glycans --fix_standalone_glycans. AlphaFold 3 was trained with leaving atoms removed even from glycans bonded to nothing; this keeps them, away from the regime it was trained in.
run_inference boolean checkbox no true Run inference Off: stop after the MSA and template step and return the completed input JSON (with the ColabFold or data-pipeline MSAs in it), for reuse in a later run.
save_embeddings boolean checkbox no false Save embeddings Write each seed's final single and pair embeddings (--save_embeddings); these are large float16 arrays.
save_distogram boolean checkbox no false Save distogram Write each seed's final distogram (--save_distogram), a num_tokens x num_tokens x 64 float16 array.
compress_large_output_files boolean checkbox no false Compress large output files Write the mmCIF and confidences JSON zstandard-compressed (--compress_large_output_files). Compressed structures are not shown in the viewer.
flash_attention_implementation string select no triton Attention implementation --flash_attention_implementation. Choose XLA for a compute capability 7.x GPU (V100, T4, RTX 20xx); the runner then also sets the XLA flag AlphaFold 3 requires for those devices. One of: triton, cudnn, xla.
jax_backend string select no gpu JAX backend --jax_backend. CPU inference needs the XLA attention implementation. One of: gpu, cpu.
gpu_device number number no 0 GPU device Zero-based JAX device index (--gpu_device), after any CUDA_VISIBLE_DEVICES filtering. minimum 0, maximum 63, step 1.
unified_memory boolean checkbox no false Unified memory Let GPU memory spill into host memory instead of failing out of memory, for GPUs with less than 80 GB: sets XLA_PYTHON_CLIENT_PREALLOCATE=false, TF_FORCE_UNIFIED_MEMORY=true and XLA_CLIENT_MEM_FRACTION=3.2, as the performance docs recommend. Slower.
buckets string text no — Compilation buckets (optional) --buckets: strictly increasing, comma-separated token counts that inputs are padded to, so a compiled model can be reused. Blank keeps AlphaFold 3's defaults (256 to 5120). up to 400 characters.

Molecule entry keys

KeyMeaning
type Which of the field's molecule types this entry is.
id A chain/entity ID, or a list of IDs for identical copies. Tools with configurable IDs expose this inside each molecule box. CatPred, whose boxes are not chains of one structure, uses it for the enzyme's sequence ID (its pdbpath, which must name exactly one sequence) and for a substrate's optional label.
count The number of identical copies. Used by tools whose native schema represents copy count separately from IDs.
sequence The residues, for a protein, dna, or rna entry.
ligand A SMILES string or a CCD_ code, for a ligand entry. CatPred resolves no CCD codes and takes SMILES only.
ion An ion code, for an ion entry.
cyclic Whether a polymer chain is cyclic. Tools that cannot model one reject it rather than ignoring it.
modifications Substitutions, as {"position": <integer>, "residue": "<CCD code>"} objects. Position indexing follows the tool: ESMFold2 uses zero-based positions; the other molecule-builder tools use one-based positions.
msa A tool-native MSA for supported protein/RNA entities. Boltz-2: an absolute .a3m/.csv path, or empty for single-sequence; omit it to generate one with the MSA server. ESMFold2: an absolute .a3m path, a serialized MSA, or a {"sequences": [...]} object.
paired_msa_path An OpenDDE, Protenix, or AlphaFold 3 protein paired-MSA path.
unpaired_msa_path An OpenDDE, Protenix, or AlphaFold 3 protein or RNA unpaired-MSA path.
templates_path An OpenDDE or Protenix protein template-hits path.
Example

A request that runs

These are the defaults, exactly as the web form would post them.

curl -X POST https://www.athanortools.com/api/alphafold3/ \
  -H 'Content-Type: application/json' \
  -d '{
  "input_mode": "parameters",
  "job_name": "af3-job",
  "sequence_molecules": "[{\"type\": \"protein\", \"id\": \"A\", \"count\": 1, \"sequence\": \"MQIFVKTLTGKTITLEVEPSDTIENVKAKIQDKEGIPPDQQRLIFAGKQLEDGRTLSDYNIQKESTLHLVLRLRGG\", \"cyclic\": false, \"modifications\": []}]",
  "model_seeds": "1",
  "bonded_atom_pairs": "",
  "user_ccd": "",
  "msa_source": "colabfold",
  "msa_server_url": "https://api.colabfold.com",
  "max_template_date": "2021-09-30",
  "resolve_msa_overlaps": true,
  "num_diffusion_samples": 5,
  "num_recycles": 10,
  "num_seeds": "",
  "conformer_max_iterations": "",
  "fix_standalone_glycans": false,
  "run_inference": true,
  "save_embeddings": false,
  "save_distogram": false,
  "compress_large_output_files": false,
  "flash_attention_implementation": "triton",
  "jax_backend": "gpu",
  "gpu_device": 0,
  "unified_memory": false,
  "buckets": ""
}'

The reply is 202 with a queued job; poll its status_url until status is succeeded or failed. See the quick start for the whole exchange.

Responses

What comes back

statusMeaning
queued Accepted, waiting for the jobs ahead of it. `position` counts how many those are.
running The tool is executing now.
succeeded Finished; `result` holds the tool's output and `license` the terms it came under.
failed Finished; `error` holds a code and a message.
cancelled Abandoned at the submitter's request; there is no result. A job cancelled before its turn never ran at all.

Errors

codeMeaning
invalid_input The client supplied invalid or incomplete input.
tool_unavailable The requested third-party dependency is not available on this host.
execution_failed A configured third-party tool exited unsuccessfully.
internal_error An adapter failed in a way it does not describe. The detail is in the server log, not the response.
not_found No job has that id. Finished jobs are dropped eventually.