StapleBridge / data /README.md
pranamanam's picture Jingjie00's picture
Upload Staplebridge files (#1)
bb6d2aa
|
Raw
History Blame Contribute Delete
1.98 kB

Data

The processed training and validation data are not included in this release. They will be handled separately. This directory is a placeholder; nothing here downloads, reconstructs, or redistributes a dataset.

Where to place files

Training reads exactly two files, resolved from the config (configs/staplebridge_main.yaml):

data:
  root: data/real          # relative paths resolve against the release root
  train_file: train.jsonl
  valid_file: valid.jsonl
  max_lead_length: 25

giving the default layout:

data/real/
├── train.jsonl
└── valid.jsonl

Any location works — set data.root to an absolute path if you prefer.

Expected input schema

JSON Lines: one object per line, loaded into staplebridge.data.schemas.LeadExample by staplebridge.data.dataset.load_leads.

Field Type Required Meaning
example_id str yes Unique lead identifier; also the property-cache key.
linear_sequence str yes Linear lead peptide, one letter per residue.
protected_positions list[int] yes (may be []) 0-based positions edits must not touch; enforced as a hard constraint.
target_context object yes in practice Must contain peptide_ca (below).
target_id str | null no Binding-partner identifier.
preferred_property_direction object no Not used by the main objective.
thresholds object no Not used by the main objective.
known_active_motif_positions list[int] | null no Optional motif annotation.

target_context.peptide_ca must be a list of [x, y, z] Cα coordinates, one per residue of linear_sequence, in order. Staple-geometry feasibility (the ca_window span check gating every candidate plan) is computed from these; a lead without them cannot yield a feasible plan support. Other keys in target_context are ignored.

Leads longer than max_lead_length are filtered out before use.