// Methodology

What actually runs

Every method, version, licence and training cutoff behind a result, and the things this platform will not tell you. The same disclosure travels with every downloadable bundle.

01 · What an analysis returns

An analysis starts from a sequence, a set of molecules to fold together, or a structure you upload. It returns structures from several independent engines, a set of checks measured on those coordinates, a per-residue account of where the engines agreed and where they did not, and the provenance of every step.

sequence / complex / uploaded structure
   -> structure prediction        OpenFold3, Boltz-2, Chai-1, independently
   -> structural checks           measured on coordinates, no model involved
   -> cross-method comparison     per residue, per engine
   -> training-set overlap        closest PDB entries and their deposition dates
   -> derived analyses            pockets, interfaces, docking, comparison
   -> bundle                      structures, JSON, provenance, parameters

The checks are reported before the comparison, deliberately. How much the methods concur is not interpretable until you know whether what they concur on contains the feature that defines your target.

02 · Structure prediction

Three engines run independently on the same input. None is treated as the reference, no consensus structure is built, and no engine is ranked above another. Where they differ, that difference is the output.

OpenFold3

0.5.0, OpenBind-0 weights · Apache-2.0 · 9 (3 model seeds × 3 diffusion samples)

Training cutoff 2025-06-30. read from the checkpoint filename of the weights actually loaded (of3-ob-2025-06-30-174k.pt) rather than a release note — the cutoff is a property of the weights, not the package version.

Boltz-2

installed in an isolated venv · MIT, code and weights · 3 diffusion samples

Training cutoff 2023-06-01. training-data section of the Boltz-2 preprint (bioRxiv 2025.06.14.659707).

Chai-1

installed in an isolated venv · Apache-2.0, code and weights · 3 diffusion samples

Training cutoff 2021-01-12. Chai-1 technical report. Whether the date means deposition or release is unresolved upstream, and we quote it as published rather than interpreting it.

OpenDDE

preview release, isolated venv · Apache-2.0, code and weights · 5 diffusion samples

Training cutoff not published. upstream publishes no training cutoff, so we do not state one. We will not infer a date to fill the column — an inferred cutoff would turn the training-overlap check below into a guess while still looking like a fact. It is therefore excluded from that check.

RhoFold+

RNA only · Apache-2.0 weights; its training data is non-commercial, so we run inference and never fine-tune · 1 (single forward pass, not a sampler)

Training cutoff not published. no cutoff published upstream; excluded from the training-overlap check for the same reason as OpenDDE.

All runs are apo by default — no template, no ligand unless you supply one.

Multiple sequence alignments

Every protein chain is folded with an alignment. It is built on our own hardware with MMseqs2 against UniRef30 (release 2023_02), a public database of 36,293,491 sequence clusters that we host rather than own — the point is where the search runs, not whose data it is. No sequence you submit leaves this infrastructure, and no third-party alignment server is contacted.

We measured what the alignment is worth, on 60 PDB targets deposited 2021–2026, pre-registered before the run: +0.165 lDDT (95% confidence interval +0.127 to +0.202).

The more useful comparison is against doing no modelling at all. For each target we took the closest relative deposited before the engine’s training cutoff and scored that structure as if a model had produced it. Copying it gives a median lDDT of 0.83–0.87, which is a high bar. Boltz-2 without an alignment was worse than copying on 41 of 54 targets; with one it was better on 32 of 51. So the alignment is the difference between output worth having and output a template beats — and the margin over copying is +0.03 to +0.05, which is smaller than a leaderboard implies.

RNA chains get an alignment too — blastn against RNAcentral (46,210,324 sequences, 37.2 billion bases), also on our own hardware. A tRNA query returns an alignment of about 17,000 sequences in under ten seconds. DNA gets none: no engine here reads an alignment for a DNA entity.

A designed RNA will often get no alignment, and that is the correct answer. Natural RNAs — riboswitches, tRNA, rRNA, viral elements — have thousands of relatives in RNAcentral. Synthetic constructs and selected aptamers frequently have none: one benchmark target we tested returned zero homologs at an e-value of 0.001. Where that happens the result says the alignment had depth 1, rather than implying one was used.

Which engine can use which alignment is narrower than which chains have one, and we checked each engine’s source rather than assuming:

  • Boltz-2, Chai-1 — protein alignments only. Chai-1 returns a single-sequence context for any non-protein entity; Boltz-2 reads the field for protein entities only.
  • OpenDDE — protein and RNA. It refuses to run without an alignment by default, so it is skipped rather than degraded when one is unavailable.
  • RhoFold+ — RNA, natively.
  • OpenFold3 — none. It accepts alignments only through its own integration, not as a file, so it runs single-sequence here.

So an RNA alignment reaches two of the five engines and is ignored by three. That is worth stating rather than implying: a comparison in which one method had more information is more useful known than hidden. Every result records, per chain, which alignment was used and how deep it was, so you never have to infer which case applied.

The RNA effect is not yet measured. The +0.165 lDDT above is a protein number from protein targets, and we are not going to quote it for RNA. Until an RNA round is run, the alignment is offered as a capability with an unmeasured effect, and the result says so.

If a search fails or times out, the run continues on the single sequence and the result says so, with the reason. It is not silently downgraded.

Engines contribute unequal numbers of structures. Each engine casts one vote per residue — its own majority across its samples — so a method that sampled more has a steadier vote than one that sampled less. Every comparison says so on the result page.

03 · Structural checks

Measured on the coordinates, with no model asked for an opinion. Base pairs are annotated by Leontis–Westhof geometry using barnaba; a G-tetrad is four guanines forming a closed cycle joined through Hoogsteen edges. Residues are keyed on a positional index rather than PDB residue numbers, so the annotation is unaffected by how a file happens to be numbered.

  • Base pairing, canonical and wobble, counted separately.
  • G-tetrads, and whether a declared G-quadruplex contains any.
  • Secondary structure in dot-bracket notation, derived from the coordinates rather than predicted from sequence.
  • Composition, and contradictions between what you declared and what the file contains.
  • Model confidence as the engine reports it, never rewritten.
  • Steric clashes, and whether an interface exists at all.

These exist because of a specific failure. A 90-compound virtual-screening benchmark in our own earlier work ran against eighteen predicted receptors that contained zero G-tetrads — on a G-quadruplex target. Every replicate agreed with every other. A model-independent check that the defining feature is physically present would have caught it before months of downstream work, and no confidence metric did.

04 · Cross-method comparison

One discrete state per residue per engine, compared position by position. The primitive depends on the molecule, because the question does:

  • Nucleic acid — base-pairing state. How did the methods fold it.
  • Protein — DSSP secondary-structure state. Helix, strand or coil.
  • Complex — contact state with a named partner chain, one comparison per partner.

Agreement and disagreement are collapsed into contiguous regions and reported as regions, with the sequence under each differing region characterised — a repeat, a homopolymer, a low-complexity tract or ordinary sequence. A low-complexity repeat has no single defined structure to predict, so methods differing across one is expected; that is a fact about the input, and it is reported as one rather than used to explain away the rest.

No consensus structure is produced, no engine is ranked, and no agreement score is emitted.

05 · Training-set overlap

The single fact that makes agreement interpretable. Query sequences are searched against PDB sequences with BLAST; deposition dates come from the wwPDB index and are compared against each engine’s published training cutoff.

Three engines agreeing on a sequence whose identical twin was deposited in 2019 is a different statement from three engines agreeing on a sequence with no close relative deposited before any cutoff. Same agreement, opposite meaning — which is why the novelty finding is bound into the comparison output rather than printed beside it.

No novelty score and no novel / not-novel verdict is computed. Only the matches, their dates, and how those sit against each cutoff.

Limitation, stated because it bit us: the routine search returns at most 25 hits, and BLAST orders by score rather than date. The earliest deposition in a capped result is therefore not necessarily the earliest that exists, and quoting it as such understates exposure. Published claims are re-run without a cap.

06 · Pocket detection

Cavities are detected with fpocket. On nucleic acids it runs with the RNA-tuned alpha-sphere and clustering parameters from Veenbaas et al. (PNAS 2025), which lift positive predictive value from 19% at protein defaults to 78% on RNA–ligand complexes. On protein it runs at its defaults, which are correct there.

fpocket’s druggability score is discarded on nucleic acids. It is a learned function trained on protein binding sites; measured against real RNA sites it consistently scores them at or near zero while rewarding cavities that are not binding sites at all. What is reported is detection and geometry — where the cavities are, which residues lie within 5 Å of each centre, and which survive across the different methods’ structures. Not a ranking of how druggable they are.

On protein the druggability score is used, because protein binding sites are what it was trained on. The same number is meaningful in one place and not the other, and the platform treats it accordingly.

07 · Ligand docking

GNINA, in a 24 Å cube centred on a detected cavity, with the box centre and size reported alongside the result. A pose search with a hidden box is not reproducible; a stated box is a statement about where we looked.

Docking is offered on protein only. On protein the search recovers near-native poses. On nucleic acids we measured it at chance for screening: it finds the site and then ranks the near-native pose last, so there is no way to say which pose a user should look at. Offering it anyway would be selling a capability we have measured the absence of.

Scores emitted by GNINA travel with the pose and are labelled as its output. CNNaffinity is a neural network’s number in pK units. It is not a measured affinity and nothing here treats it as one.

08 · Interfaces

For a complex: which residues contact which partner, per method, with steric clashes reported as checks rather than as a footnote. Molecules are identified by sequence rather than by chain letter, because engines rename chains freely and matching on the letter silently compares one engine’s RNA against another’s protein. Contacts are reported at alignment positions, not raw residue numbers, for the same reason.

09 · Reproducibility

Every result carries the method, its version, the parameters, the database versions and the timings, and the bundle is downloadable. Reproducibility differs by engine and we state it per engine rather than averaging over the difference:

  • Boltz-2 and Chai-1 accept a seed value. At a fixed seed Boltz-2 returns bit-identical structures across runs, verified by hashing the output.
  • OpenFold3 exposes a seed count and no seed value. Its individual structures are not reproducible run to run; its sequence-level results are.

Third-party tools that need a different Python or numpy version than the main environment run in their own virtual environments as subprocesses, communicating over JSON. That keeps versions pinned independently, and it keeps copyleft-licensed tools out of this codebase — see below.

10 · What we do not claim

  • No composite score, no target-quality score, no go/no-go verdict. “Validated against what?” has no answer for a target score, and every composite we built was killed by its own controls.
  • Agreement between methods is not evidence. Engines trained on overlapping snapshots of the same database agree for reasons that have nothing to do with correctness. Read it alongside the training-set overlap.
  • We do not know whether a molecule binds. Detection of a cavity is not a claim about affinity, and no metric we have tested separates real binders from decoys on RNA.
  • A fold class is never declared out of scope without naming the engine it was measured on. Capability belongs to the pair (engine, fold class). On a telomeric G-quadruplex one engine produced the correct tetrad topology and two produced none; on a four-way junction all three were perfect. The two behave nothing alike.
  • Detection is not accuracy. That a feature is present does not mean the structure is correct in any other respect.

11 · Licences and isolation

Third-party tools run server-side only. Two are copyleft-licensed and are invoked as separate processes over files, never imported, linked or shipped in anything a customer receives.

OpenFold3 0.5.0Apache-2.0
Boltz-2MIT (code and weights)
Chai-1Apache-2.0 (code and weights)
fpocketMIT
barnaba 0.1.9GPL-3.0 — isolated venv, subprocess, JSON over stdout
GNINAGPL (via Open Babel) — isolated subprocess, protein only

barnaba is isolated because it needs an older numpy than the main environment. The licence isolation is a consequence of that rather than its purpose, but it is real and worth stating: no GPL code is linked into this platform.

12 · Citations

  • Abramson, J. et al. (2024).Accurate structure prediction of biomolecular interactions with AlphaFold 3.Nature 630, 493–500.
  • Passaro, S. et al. (2025).Boltz-2: towards accurate and efficient binding affinity prediction.bioRxiv 2025.06.14.659707.
  • Chai Discovery team (2024).Chai-1: decoding the molecular interactions of life.bioRxiv.
  • Bottaro, S., Bussi, G., Pinamonti, G., Reißer, S., Boomsma, W., Lindorff-Larsen, K. (2019).Barnaba: software for analysis of nucleic acid structures and trajectories.RNA 25, 219–231.
  • Leontis, N.B., Westhof, E. (2001).Geometric nomenclature and classification of RNA base pairs.RNA 7, 499–512.
  • Le Guilloux, V., Schmidtke, P., Tuffery, P. (2009).Fpocket: an open source platform for ligand pocket detection.BMC Bioinformatics 10, 168.
  • Veenbaas, S.D., Koehn, J.T., Irving, P.S., Lama, N.N., Weeks, K.M. (2025).Ligand-binding pockets in RNA and where to find them.PNAS 122(17), e2422346122.
  • McNutt, A.T. et al. (2021).GNINA 1.0: molecular docking with deep learning.Journal of Cheminformatics 13, 43.

// Superseded

An earlier version of this page documented the v0.2 pipeline: RhoFold+, an ANM conformational ensemble, and a ranked shortlist of candidate druggable pockets. That is no longer what runs, and two of its central claims were later disproved by our own controls. It is kept, labelled, with what was measured against it.

Read the superseded v0.2 methodology →