Description
SimpleFold-Turbo: Adaptive Inference Caching Yields 14-fold Acceleration of Flow-matching Protein Structure PredictionGeneral Information for Raw DataDescription: This dataset contains all benchmarking data, predicted structures, and analysis results for the SimpleFold-Turbo manuscript. SimpleFold-Turbo applies TeaCache-style adaptive step-skipping to SimpleFold diffusion models across six model scales (100M–3B parameters), evaluated on a structurally diverse subset of 300 CATH domains.Total size: ~572 MB (553 MB predicted structure files)Benchmark SetFileDescriptionCATH300.csvBenchmark set of 300 CATH domains: name, sequence, and lengthdiverse_cath_300.fastaSequences for the 300 domains in FASTA formatdiverse_cath_300.jsonExtended metadata for each domain (CATH classification, structural annotationsss_content.jsonSecondary structure content (helix/sheet/coil fractions) per domainCore Benchmark ResultsFileDescriptioncath_benchmark_full.csvPer-protein benchmark results (3,600 rows): model, TeaCache threshold, TM-score, RMSD, lDDT, inference time, and cache hit ratecath_benchmark_full.jsonSame data in JSON formatgt_comparison.csvSide-by-side quality comparison (baseline vs. TeaCache) against ground-truth experimental structures: TM-score, RMSD, and lDDT differencesDual Sweep (Uniform vs. Adaptive Step-Skipping)FileDescriptiondual_sweep_simplefold_{100M,360M,700M,1.1B,1.6B,3B}.jsonPer-protein results for each model scale (5,100 entries each). Each entry records method (uniform or adaptive), condition (number of steps or threshold), inference time, cache hit rate, computed steps, and quality metrics (RMSD, TM-score)dual_sweep_summary.csvAggregated summary across all models and conditions (30,600 rows)uniform_vs_adaptive.csvHead-to-head comparison of uniform vs. adaptive skipping at matched compute budgetsThreshold and Step SweepsFileDescriptionthreshold_sweep.jsonPer-protein results across TeaCache threshold values (2,400 entries)threshold_summary.csvAggregated: mean time, speedup, cache hit rate, and quality loss per thresholduniform_step_sweep.jsonPer-protein results for uniform step counts (2,700 entries)Mechanistic AnalysesFileDescriptionskip_patterns.jsonTimestep-resolved skip/compute patterns across the denoising trajectory. Includes per-step skip rates, a summary of always-computed warmup steps (11) vs. always-skipped steps (200) vs. variable steps (289)warmup_comparison.jsonAnalysis of warmup phase: compares the first 11 (always-computed) steps to full 500-step trajectories across 300 proteinsclustering_results.jsonClustering of denoising timesteps into two regimes based on skip behavior, with secondary-structure correlationcrystallization_results.jsonAtom-level settling ("crystallization") analysis: per-protein statistics on when atomic coordinates stabilize during denoising (20 proteins)dimensionality_control.csvCache hit rate vs. chain length and embedding dimensionality (synthetic and empirical)dimensionality_control.jsonFull dimensionality control experiment data including Pearson correlationsPredicted Structuresstructures.zip:30,811 PDB files (~553 MB compressed). Organized as structures/{model}/{method}{condition}/{domain}.pdb, where model is one of simplefold{100M,360M,700M,1.1B,1.6B,3B}, method is uniform or adaptive, and condition is the step count or threshold value.File Formats- CSV files use comma delimiters with a header row- JSON files are either arrays of per-protein result objects or dictionaries with descriptive top-level keys- FASTA follows standard format with CATH domain identifiers as headers- PDB files follow standard Protein Data Bank formatReproducing the FiguresThe Python scripts used to generate all manuscript figures from these data files are included in the GitHub repo publication/ directory (figure1.py, figure2.py, figure_supplement.py).
Citations (0)
No citations found
Mentions (0)
No mentions found
Metrics Over Time
Publication Details
DOI
Publisher
Zenodo
Subfield
Civil and Structural Engineering
Field
Engineering
Domain
Physical Sciences
Confidence Score
42%
Source
Scholar Data Model