Raw data for SimpleFold-Turbo preprint

Taghon, Geoffrey

Description

SimpleFold-Turbo: Adaptive Inference Caching Yields 14-fold Acceleration of Flow-matching Protein Structure PredictionGeneral Information for Raw DataDescription: This dataset contains all benchmarking data, predicted structures, and analysis results for the SimpleFold-Turbo manuscript. SimpleFold-Turbo applies TeaCache-style adaptive step-skipping to SimpleFold diffusion models across six model scales (100M–3B parameters), evaluated on a structurally diverse subset of 300 CATH domains.Total size: ~572 MB (553 MB predicted structure files)Benchmark SetFileDescriptionCATH300.csvBenchmark set of 300 CATH domains: name, sequence, and lengthdiverse_cath_300.fastaSequences for the 300 domains in FASTA formatdiverse_cath_300.jsonExtended metadata for each domain (CATH classification, structural annotationsss_content.jsonSecondary structure content (helix/sheet/coil fractions) per domainCore Benchmark ResultsFileDescriptioncath_benchmark_full.csvPer-protein benchmark results (3,600 rows): model, TeaCache threshold, TM-score, RMSD, lDDT, inference time, and cache hit ratecath_benchmark_full.jsonSame data in JSON formatgt_comparison.csvSide-by-side quality comparison (baseline vs. TeaCache) against ground-truth experimental structures: TM-score, RMSD, and lDDT differencesDual Sweep (Uniform vs. Adaptive Step-Skipping)FileDescriptiondual_sweep_simplefold_{100M,360M,700M,1.1B,1.6B,3B}.jsonPer-protein results for each model scale (5,100 entries each). Each entry records method (uniform or adaptive), condition (number of steps or threshold), inference time, cache hit rate, computed steps, and quality metrics (RMSD, TM-score)dual_sweep_summary.csvAggregated summary across all models and conditions (30,600 rows)uniform_vs_adaptive.csvHead-to-head comparison of uniform vs. adaptive skipping at matched compute budgetsThreshold and Step SweepsFileDescriptionthreshold_sweep.jsonPer-protein results across TeaCache threshold values (2,400 entries)threshold_summary.csvAggregated: mean time, speedup, cache hit rate, and quality loss per thresholduniform_step_sweep.jsonPer-protein results for uniform step counts (2,700 entries)Mechanistic AnalysesFileDescriptionskip_patterns.jsonTimestep-resolved skip/compute patterns across the denoising trajectory. Includes per-step skip rates, a summary of always-computed warmup steps (11) vs. always-skipped steps (200) vs. variable steps (289)warmup_comparison.jsonAnalysis of warmup phase: compares the first 11 (always-computed) steps to full 500-step trajectories across 300 proteinsclustering_results.jsonClustering of denoising timesteps into two regimes based on skip behavior, with secondary-structure correlationcrystallization_results.jsonAtom-level settling ("crystallization") analysis: per-protein statistics on when atomic coordinates stabilize during denoising (20 proteins)dimensionality_control.csvCache hit rate vs. chain length and embedding dimensionality (synthetic and empirical)dimensionality_control.jsonFull dimensionality control experiment data including Pearson correlationsPredicted Structuresstructures.zip:30,811 PDB files (~553 MB compressed). Organized as structures/{model}/{method}{condition}/{domain}.pdb, where model is one of simplefold{100M,360M,700M,1.1B,1.6B,3B}, method is uniform or adaptive, and condition is the step count or threshold value.File Formats- CSV files use comma delimiters with a header row- JSON files are either arrays of per-protein result objects or dictionaries with descriptive top-level keys- FASTA follows standard format with CATH domain identifiers as headers- PDB files follow standard Protein Data Bank formatReproducing the FiguresThe Python scripts used to generate all manuscript figures from these data files are included in the GitHub repo publication/ directory (figure1.py, figure2.py, figure_supplement.py).

Citations (0)

Mentions (0)

Metrics

Dataset Index

0.8

FAIR Score

88%

Citations

1

Mentions

0

Metrics Over Time

Publication Details

DOI

Publisher

Zenodo

License

Creative Commons Public Domain Dedication and Certification

Copyright (C) 2026 Department of Commerce, United States of America

Assigned Domain

Subfield

Civil and Structural Engineering

Field

Engineering

Domain

Physical Sciences

Confidence Score

42%

Source

Scholar Data Model

Keywords

Protein Structure, TertiaryMachine Learning

Normalization Factors

FT

65.38

CTw

1.00

MTw

1.00