Each feature is tagged by where its data comes from: deposited metadata taken verbatim from the PDB entry · tool-derived produced from the deposited coordinates by an established tool (x3dna-dssr, US-align, BLAST, RDKit) · geometry-derived measured directly from the coordinates here (gemmi / NumPy) · heuristic a rule-based inference — treat as a hypothesis.
holo_len / apo_len columns count
modeled nucleotides. A served pair can therefore differ by more than 20% there without
failing this gate (8OLZ vs 9QU7: 23.8% modeled but 1.3% declared — 9QU7 models 297 of its
396 nt).Net effect on dataset bias: a ≥90% apo–holo identity floor, same-pocket and ≤20%-length-difference matching, and an RNA-only restriction — a curated subset served out of the larger underlying set. The Database gallery opens on the full release - both identity facets start on the whole span present in the data, so nothing is hidden until you narrow one. The bulk download always contains the full served set.
| Feature | What you see here | Where it comes from |
|---|---|---|
| Basic Information | both structures' method/resolution/authors + ligand properties | RCSB Data API (live) + RDKit deposited metadatageometry-derived |
| 3D superposition + morph | the superposed apo and fragment-bound holo aptamer | sequence-anchored Kabsch on the paired C1′ atoms (a US-align fit is computed alongside and the lower-RMSD one kept; Kabsch is the analytic minimiser of that RMSD, so it is what wins on every served record); morph = interpolation, not MD tool-derivedheuristic |
| Per-residue displacement | which nucleotides move, and how far (C1′) | sequence-anchored Kabsch (lowest-RMSD transform; see above) tool-derived |
| Conformational dashboard (feature track) | a per-nucleotide apo→holo C1′ displacement heatmap over the holo sequence (with a x3dna-dssr motif band); radius of gyration and ligand burial are reported as static apo/holo endpoint values in the Overview, not morph curves. A visualization between two endpoint structures, not MD | geometry on the best-of US-align / sequence-anchored Kabsch superposition (burial via biopython SASA) geometry-derived |
| Feature track (per-nucleotide) | each holo nucleotide as a cell in the feature track: apo→holo displacement color + binding-site marker + DSSR-motif band; click → 3D | geometry on displacement + interactions + DSSR motifs geometry-derived |
| TM-score & radius of gyration (overall fold) | overall fold similarity (TM-score) and global compaction (Rg, apo → holo) on the Overview. Rg is measured over the apo–holo shared (aligned-core) residues. TM-score is length-sensitive: below ~30 nt it tracks chain length, so read high = fold conserved (not low = fold change) for short RNA | US-align TM-score + gemmi Rg over shared-core C1′ geometry-derived |
| Sequence alignment (apo vs holo) | the two states lined up residue-by-residue | Needleman–Wunsch geometry-derived |
| Secondary structure (2D) | base-pairing of apo vs holo; click a base → 3D - what DSSR extracts ↓ | x3dna-dssr tool-derived |
| Ligand–RNA interactions (3D) | H-bonds / contacts to the fragment, with Å labels | gemmi contact geometry on the holo structure; H-bond vs contact typing from x3dna-dssr
--get-hbond geometry-derivedtool-derived (x3dna-dssr) |
| Interaction-data tables | pairs · H-bonds · water bridges · base pairs · BP changes · motifs - what DSSR extracts ↓ | DSSR (pairs/motifs) + gemmi (contacts) tool-derivedgeometry-derived |
| Statistics (site-wide) | where this pair sits among all pairs (Δ, resolution, composition) | aggregation geometry-derived |
the sequence and dot-bracket notation of the folded chain → drives the Secondary structure (2D) diagram.
every pair - canonical and non-canonical - with the bases (e.g. G–U), the Leontis–Westhof geometry type (e.g. cWW, tHS) and the Saenger class → the Base pairing table; the apo–holo set difference becomes Base-pair rewiring.
stems, hairpin loops, internal loops, bulges, multi-way junctions and single-strand segments (type · residues · size) → the Structural motifs table.
Each base has three edges - Watson–Crick, Hoogsteen, and Sugar. A base pair is named by the glycosidic-bond orientation (cis or trans) plus the two interacting edges — 12 geometric families in all. Examples: cWW = cis Watson–Crick/Watson–Crick (the standard A–U, G–C, G–U pairs); tHS = trans Hoogsteen/Sugar-edge (a common non-canonical pair). This geometry is read from the experimental 3D structure, so non-canonical pairs the dot-bracket can't show still appear in the Base pairing table.
Underlying structures come from RCSB PDB; the pair list & pockets from the SMARTFlexDB apo–holo dataset. Everything else is computed locally by x3dna-dssr, US-align, RDKit and gemmi. Note: the morph is a straight-line interpolation between the two endpoint structures (not a molecular-dynamics trajectory). The RNA type is NAKB's curated functional class of the host chain (with a title-keyword fallback for the few structures NAKB does not annotate).
The SMARTFlexDB derived dataset — the apo–holo pair list, geometric descriptors, alignments and cluster assignments — is released under CC‑BY‑4.0. The underlying atomic coordinates are retrieved from the RCSB PDB and remain subject to the PDB's terms of use — no experimental data are redistributed beyond what is already public in the PDB.
SMARTFlexDB is built on these tools: x3dna-DSSR (secondary structure & motifs); US-align (the TM-score, and a comparison fit) and a NumPy Kabsch fit (the served apo–holo superposition / RMSD); BLAST (sequence identity). It also uses MMseqs2 and RM-align (sequence / structure clustering); PLIP (interactions); RDKit (ligand properties); and gemmi (structure handling).
Questions, corrections, bug reports, or a candidate apo–holo pair we missed? Email yanjun.li@ufl.edu and jiang.shiyu@ufl.edu - we welcome feedback and respond to data corrections promptly. You can also run the full analysis on your own apo and holo structures from the Query page.