rules_lora
Bazel-native LoRA fine-tuning.
| Latest | 0.1.4 |
|---|---|
| Versions | 39 |
| Category | Bazel rules |
| Maintainers | Matt Marshall |
| Registry | https://registry.tbzl.dev/modules/rules_lora/ |
| Source | github.com/tomato-bazel/rules_lora |
bazel_dep(name = "rules_lora", version = "0.1.4")
View source & releases on GitHub ↗
Bazel-native LoRA fine-tuning — hermetic, reproducible low-rank-adapter training as Bazel targets.
Use it
# MODULE.bazel — resolves from the fastverk registry (registry.fastverk.com)
bazel_dep(name = "rules_lora", version = "0.0.35")
See the package BUILD.bazel / defs.bzl for the training + adapter rules.
Part of the tomato-bazel distribution.
Usage#
Real usage, taken from the module’s examples/.
examples/corpus_smoke/BUILD.bazel
load("@rules_python//python:defs.bzl", "py_binary")
load(
"@rules_lora//lora:defs.bzl",
"lora_corpus",
"lora_dataset",
)
# Smoke example for `lora_corpus` (v0.0.3).
#
# Two corpora chained: `base_corpus` transforms a raw text file into
# messages_v1 SFT rows; `derived_corpus` takes `base_corpus` as a dep
# AND its own raw source, exercising both `source` and `deps`.
py_binary(
name = "smoke_transform",
srcs = ["transform.py"],
main = "transform.py",
)
lora_corpus(
name = "base_corpus",
source = ["raw_prompts.txt"],
transform = ":smoke_transform",
schema = "messages_v1",
min_examples = 3,
)
lora_corpus(
name = "derived_corpus",
source = ["raw_prompts.txt"],
transform = ":smoke_transform",
deps = [":base_corpus"],
schema = "messages_v1",
# 3 (source) + 3 (dep) = 6 rows expected.
min_examples = 6,
)
# Demonstrate that a `lora_corpus` plugs in wherever `lora_dataset`
# is accepted: alias the corpus as a dataset for downstream training
# pipelines.
alias(
name = "derived_corpus_as_dataset",
actual = ":derived_corpus",
)
examples/smoke/BUILD.bazel
load(
"@rules_lora//lora:defs.bzl",
"lora_base_model",
"lora_dataset",
"lora_lineage",
"lora_merge",
"lora_recipe",
"lora_train",
)
# Smoke test exercising all four public macros end-to-end.
# `bazel build //examples/smoke:smoke_jobspec` produces a
# TrainingJobSpec JSON; no actual training happens (the `run`
# subcommand of runpod_orchestrator is v0.1).
lora_dataset(
name = "tiny_sft",
src = "sft.jsonl",
min_examples = 3,
schema = "messages_v1",
)
lora_recipe(
name = "tiny_recipe",
alpha = 8,
epochs = 1,
framework = "torchtune",
grad_accum_steps = 1,
learning_rate = "5e-4",
micro_batch_size = 1,
rank = 4,
target_modules = [
"q_proj",
"v_proj",
],
)
lora_base_model(
name = "tiny_base",
repo = "Qwen/Qwen2.5-0.5B-Instruct",
revision = "main",
)
lora_train(
name = "smoke_jobspec",
base = ":tiny_base",
dataset = ":tiny_sft",
recipe = ":tiny_recipe",
# Backend is per-platform now: defaults to runpod; build with
# --platforms=@rules_lora//lora/backend:local_platform for the local path.
)
# Exercises the lora_merge macro (rule + .run wrapper). No execution
# here — `bazel build //examples/smoke:smoke_merge` just checks the
# rule analyzes and the wrapper script is generated.
lora_merge(
name = "smoke_merge",
adapter_dir = "outputs/adapter-smoke_jobspec",
base = ":tiny_base",
out_dir = "outputs/merged-smoke",
# No push_repo here — merge-only, so the smoke build stays
# side-effect-free. Set push_repo + `bazel run :smoke_merge.run`
# to also push the merged dir to the hub.
)
# Provenance: trace smoke_jobspec's transitive dataset/recipe/base lineage.
# `bazel build //examples/smoke:smoke_lineage` -> smoke_lineage.lineage.json.
lora_lineage(
name = "smoke_lineage",
target = ":smoke_jobspec",
)Rules & providers#
Generated with Stardoc from the module's .bzl sources.
from docs/hermetic-runners-roadmap.md
Hermetic LoRA runners — roadmap
Status: step 1 (lock the deps) DONE + real-hardware verified; steps 2–5
(Bazel wheel-vendoring) deliberately banked. Both runners now pin their dep
sets — local to the MPS regime (fresh-venv verified), runpod to a torch-2.4 set
matched to its base image (verified on a real RunPod CUDA pod). That removes the
fragility that motivated this doc. Vendoring the wheels via pip.parse is an
optimization of working, reproducible paths — and the training itself is
inherently non-hermetic (it needs the GPU/MPS host and downloads the base
model), so the marginal value is low. Decision (2026-06): do not build the
vendoring rewrite until there’s a concrete driver (e.g. offline/air-gapped CI,
or a hard no-runtime-network requirement). The steps below stay as the plan for
that day. A hollow scaffold (stubbed pod lifecycle, fake torch lock, half-wired
select()) is worse than the working pinned venv + @rules_runpod paths that
ship today — don’t merge one.
Where we are now (shipped)
The backend dispatch + de-shell work already landed (PR #1 on main):
- Per-platform backend toolchain (
//lora/backend):local/runpod/modalare registered toolchains selected by--platforms;lora_trainresolves the toolchain (LoraBackendInfo) for the jobspec composer + backend identity. The.runentry dispatches on the same:backendconstraint viaselect(). - No generated shell of ours. The local + merge run entries are
py_binaryorchestrators reading build-generated JSON configs (*.local.json,*.merge.json);local_runner.shand the bash wrappers are gone. - What is still non-hermetic (by design, deferred here):
runtime/local_runner/local_train.pyis a thin orchestrator — it creates a runtime venv,pip installs a now fully-pinned dep set (the validated MPS regime), downloads the HF model, and shellstune run. Reproducible, but still non-hermetic: runtime network + host accelerator, not Bazel-vendored.runtime/runpod_orchestrator/src/main.rsdoes build-time manifest synth only (write-jobspec+write-runpod-manifest); the deadrunstub was removed (it duplicated@rules_runpod). The working RunPod path is the@rules_runpodmacro composition.
Goal
Two independent tracks, each ending in a runner binary the backend toolchain
points at, with no runtime venv and no @rules_runpod dependency.
Track 1 — hermetic local runner (vendored torch)
Replace the runtime venv with Bazel-vendored deps + in-process torchtune.
- Lock the training deps. ✅ DONE (both runners, real-hardware verified).
The local runner’s
_PIP_PINSis fully pinned to the validated MPS regime (torch==2.12.0,torchtune==0.4.0,torchao==0.5.0,kagglehub==0.2.9,huggingface_hub[cli]==1.19.0,transformers==5.12.0,datasets==5.0.0), verified by a fresh-venv train on Apple-Silicon MPS. The runpod backend’s setup is pinned to a separate torch-2.4 set matched to its base image (rules_lora 0.1.2), verified on a real RunPod CUDA pod. The two regimes can’t share one set (different torch). Landmine remains: the triangle breaks on unpinned releases — treat any bump as a deliberate, tested change on the target Python. What’s not done is steps 2–5: turning these pinned sets into Bazel-vendored wheels (they’re still runtimepip installs). pip.parseinMODULE.bazelover the lock, exposing@lora_pip//torch,@lora_pip//torchtune, etc. Use platform-conditional requirement sets: MPS (mac arm64) vs CPU vs CUDA wheels are different downloads —select()the right@lora_pip_{mps,cpu,cu121}//...per--platforms. This is the part that needs care; analysis can check the labels resolve, only a real fetch + import confirms the wheel set.- HF base model as a Bazel artifact. A repository rule (or
http_file/http_archive) that fetches the pinnedbase_id@revisionsnapshot into a repo, so the model is an input, not a runtimehf download. Large; cache-aware. (Or keep the runtime download as the explicit “non-hermetic edge” and document it — vendoring multi-GB models in Bazel is a real cost/benefit call.) - In-process torchtune. Rewrite
local_train.pytoimport torchtuneand invokelora_finetune_single_devicevia its Python API against the rendered config — notune runCLI subprocess, no venv activation. Thepy_binarythendepson the vendored torch/torchtune instead of building a venv. - Wire it into the toolchain. The
localbackend toolchain’s runner becomes this hermeticpy_binary; drop the venv path fromlocal_train.py.
Verification (needs hardware): bazel run the local backend on an
Apple-Silicon box, confirm a real LoRA step trains end-to-end (device mps),
adapter lands in outputs/adapter-<name>. Repeat on a CUDA Linux box for the
cu121 wheel set.
Baseline validated (2026-06, Apple-Silicon MPS): the current venv local
runner trains a real LoRA end-to-end (Qwen2.5-0.5B-Instruct, device mps,
adapter written to outputs/adapter-<name>), and lora_merge then folds it
into the base → a standalone, loadable HF dir (candle/CPU, 48 projections,
scale α/r). So Track 1 (hermetic vendoring) is an optimization of a working
path, not a fix. The size-mismatch builder bug found en route — the runner
hardcoded the 1.5B builder for the whole qwen2 family — is fixed in 0.1.1
(size-aware _model_builder, deriving the builder from the parsed base size).
Track 2 — RunPod backend: already implemented by rules_runpod (not a rewrite)
Correction (2026-06): the original “reimplement the pod lifecycle” framing was
wrong — checked against the actual code with a live key. rules_runpod’s CLI
already implements the full lifecycle — deploy → upload (S3 volume or SSH
rsync) → ssh → tune run → poll → download adapter → terminate — via a dedicated
runpod SDK crate (runpod::Client, REST API; see cli/src/pod.rs,
train.rs), and @rules_runpod’s runpod_job macro already drives it. The
current lora_train runpod backend works through that. So there is no
from-scratch orchestrator to build; reimplementing it in
runtime/runpod_orchestrator would just duplicate rules_runpod.
What’s actually left for the runpod backend is small and optional:
- (Optional) single-binary wiring. If you want the per-platform
runpodtoolchain runner to be one binary instead of the@rules_runpodmacro composition, haveruntime/runpod_orchestrator’srunsubcommand call therunpodcrate (the onerules_runpodalready uses) rather than reimplement the REST calls — lifecycle stays inrules_runpod, the lora side just reads the jobspec and hands off. Wiring/ergonomics, not new capability; the venv-free win is marginal for runpod (the heavy work runs on the pod anyway). - Done: the
runpod_orchestrator runstub has been deleted — it advertised a capabilityrules_runpodalready provides. The orchestrator binary now exposes only its two build-time synth subcommands.
Validation note (from a live key): RunPod’s GraphQL pod-creation is
deprecated — read queries work, the create mutation 403s on a read-scoped key;
rules_runpod correctly uses the REST API via the runpod crate. A live
training check therefore runs through the existing rules_runpod path
(write-scoped key + SSH key + a synth manifest), not a hand-rolled API call. Key
- account confirmed working (read); 44 GPU types available.
Sequencing & guardrails
- Track 1 is the real remaining work (Track 2 turned out to be mostly “delete the stub / optional wiring” — see the correction above). Do Track 1 on an Apple-Silicon box: lock the torch set → split wheels → in-process torchtune → verify a real MPS step.
- Keep the current working paths (venv local /
@rules_runpod) in place until each replacement is verified on hardware — flip the toolchain runner only when green. - Per the repo’s DTO convention, the runner↔orchestrator contract is already a
proto (
lora.v1.TrainingJobSpec); keep new config on it.
Related deferred items (not this doc’s scope, noted for completeness)
- rules_postgres Gate 3 runs only where the private
//crates/pipelineclang/LLVM tools exist (the public@rules_lang//rules/cships rule defs only). Making Gate 3 public would mean porting that ~8k-LOC Rust subsystem + clang toolchain into public rules_lang. - Stranded
atlas-v0.3.0tag on the Syntax-less polyglot commit (protected, can’t delete; no release attached, unused). The live atlas isatlas-v0.3.1viarules_lang 0.3.0.
Conformance#
5 findings across 2 invariants. 10 contested atoms. See how gating works or the full report.
| version | toolchain |
|---|---|
0.1.4 | @rules_lora//lora/backend:local_toolchain |
0.1.4 | @rules_lora//lora/backend:modal_toolchain |
0.1.4 | @rules_lora//lora/backend:runpod_toolchain |
0.1.4 | @rust_toolchains//:all |
| repo | extension |
|---|---|
crates | @rules_rust//crate_universe:extension.bzl |
Contested atoms
Third-party modules where this module resolves a different version than others do. Not a violation of anything this module did — it is the actionable form of a registry-level convergence finding, and the sentence a maintainer can act on.
| Atom | Resolved here | Elsewhere |
|---|---|---|
apple_support | 1.24.2 | 2.2.0 ×1 |
bazel_skylib | 1.8.2 | 1.9.0 ×2 |
gazelle | 0.36.0 | 0.30.0 ×50.44.0 ×10.51.0 ×3 |
nlohmann_json | 3.6.1 | 3.12.0.bcr.1 ×1 |
protobuf | 33.4 | 34.0.bcr.1 ×2 |
rules_go | 0.60.0 | 0.39.1 ×5 |
rules_jvm_external | 6.7 | 6.8 ×4 |
rules_python | 1.7.0 | 2.0.1 ×1 |
rules_swift | 3.1.2 | 3.6.1 ×1 |
upb | 0.0.0-20220923-a547704 | 0.0.0-20230516-61a97ef ×1 |
Dependencies#
Depends on
Used by (1 in the registry)
Versions#
39 published versions, newest first. Each resolves to an immutable, integrity-checked archive.
| Version | Integrity (sha256) | Source archive |
|---|---|---|
0.1.4 latest | I5j9aQPH98erPTkO… | tag archive ↗ |
0.1.3 | KoEJJCAp2mB5Lc2l… | tag archive ↗ |
0.1.2 | P6rJ/tpog3XBV0ku… | tag archive ↗ |
0.1.1 | fBkNSR3amUjLgiBp… | tag archive ↗ |
0.1.0 | 0t0J3wexszQZ5shb… | tag archive ↗ |
0.0.35 | BrzRCAhYaate/By6… | tag archive ↗ |
0.0.34 | tdfLGmX7CW09HVoO… | tag archive ↗ |
0.0.33 | xHJnJAj00TqsaTsr… | tag archive ↗ |
0.0.32 | +EPrEPILDnDlUYI2… | tag archive ↗ |
0.0.31 | qLeW93E6I+3USa3T… | tag archive ↗ |
0.0.30 | 10eBAkxbs1r5EFV8… | tag archive ↗ |
0.0.29 | ygqjKaDFTdD08Jmd… | tag archive ↗ |
0.0.28 | nSLmIi4zHkbfQcsZ… | tag archive ↗ |
0.0.27 | R6MsTwjkf6hTAiy5… | tag archive ↗ |
0.0.26 | YPI6d13pdvOn23I0… | tag archive ↗ |
0.0.25 | /DSQI8FwvcqYQMCT… | tag archive ↗ |
0.0.24 | 0agjt2nCaj+/LXG/… | tag archive ↗ |
0.0.23 | H54kupfDTLQK849K… | tag archive ↗ |
0.0.22 | Pmi+SbUIa0cnxUnq… | tag archive ↗ |
0.0.21 | fKzqtrXmkdVQbHHd… | tag archive ↗ |
0.0.20 | Y02Hr1lJ3Kcx+63p… | tag archive ↗ |
0.0.19 | VV03lrXWlbertipN… | tag archive ↗ |
0.0.18 | RlDQLh8LfbnmMOtF… | tag archive ↗ |
0.0.17 | r/ZEU75JhiFCCKlA… | tag archive ↗ |
0.0.16 | VWv0sptWnm94yWRL… | tag archive ↗ |
0.0.15 | 0s4F6vXq2EwbU4iS… | tag archive ↗ |
0.0.14 | 0YVJDVczBTSmEuaH… | tag archive ↗ |
0.0.13 | 6KygxaI6qZt1whRL… | tag archive ↗ |
0.0.12 | sWHCqR63ZTZsGzQh… | tag archive ↗ |
0.0.11 | l4mDHMW41m2BMF4X… | tag archive ↗ |
0.0.10 | qovALsEknn7iAPJy… | tag archive ↗ |
0.0.9 | fR/eJGs1QsurdqGB… | tag archive ↗ |
0.0.8 | 2zmhecnZWJp+suSs… | tag archive ↗ |
0.0.7 | 51eaMJp8uTgERMRd… | tag archive ↗ |
0.0.6 | WnikXemb/Fky3r+z… | tag archive ↗ |
0.0.5 | i8i4zO+ykgaRNhTT… | tag archive ↗ |
0.0.4 | Rx7au8Q540+31fhY… | tag archive ↗ |
0.0.3 | +6gkpgH950rOPOCP… | tag archive ↗ |
0.0.1 | IlLZ0EQwqVnjcDJV… | tag archive ↗ |
Changelog#
All notable changes to rules_lora. The format is loosely
Keep a Changelog — version headers
mirror the published bazel-registry entries. This repo is public
(github.com/fastverk/rules_lora).
0.1.3 — local runner: fully-pinned dep set (fresh-venv verified)
Applies the same fragility fix as 0.1.2 (runpod), now to the local backend.
local_train.py’s _PIP_PINS left torch/huggingface_hub/transformers/
datasets unpinned, so a fresh venv drifted onto whatever was latest — the
exact drift that broke the runpod setup. Pinned to the validated MPS regime
(torch==2.12.0, torchao==0.5.0, torchtune==0.4.0, kagglehub==0.2.9,
huggingface_hub[cli]==1.19.0, transformers==5.12.0, datasets==5.0.0) and
verified by a fresh-venv train on Apple-Silicon MPS (5 steps, real adapter).
This is a different set from the runpod backend’s torch-2.4 pins — the two
can’t share one (the local Mac uses the latest torch; the runpod path is pinned
to its older base image). Completes “lock the training deps” for both runners
(docs/hermetic-runners-roadmap.md Track 1, step 1); Bazel-vendoring the wheels
(steps 2–5) remains a dedicated, multi-platform effort.
0.1.2 — runpod backend: size-aware builder + pinned setup (real-GPU verified)
Makes the runpod backend actually train, verified end-to-end on a real
RunPod CUDA pod (RTX A4000): deploy → SSH → rsync → setup → tune run on
device: cuda → adapter pulled back. Two fixes in the manifest synthesizer
(runpod_orchestrator write-runpod-manifest):
- Size-aware torchtune builder. The synthesizer hardcoded
lora_qwen2_1_5bfor the whole qwen2 family — the same bug 0.1.1 fixed in the local runner, in the runpod code path. Any non-1.5B qwen2 base (the Qwen2.5-0.5B smoke included) failed at checkpoint load with a tensor size mismatch on the pod. The builder now derives from the parsedbase_idsize (lora_builder/parse_param_size, + arust_test). - Pinned setup dependency stack. The setup block installed unpinned
transformers/datasets/huggingface_hubon top of the image’s torch 2.4.1+cu124, pulling releases that need a newer torch (transformers 5.x) — the run failed before training. Pinned to the torch-2.4 era (transformers==4.46.3,datasets==3.1.0,huggingface_hub==0.26.2,kagglehub==0.2.9) and switchedhf download→huggingface-cli download(thehfcommand needs hf-hub ≥ 0.34). torch itself is left untouched (reinstalling it pulls a wheel built for a CUDA newer than the pod’s driver).
0.1.1 — local runner: size-aware torchtune model builder
Patch fix for the local backend. local_train.py hardcoded the
lora_qwen2_1_5b model builder for the whole qwen2 family, so any non-1.5B
qwen2 base (Qwen2.5-0.5B/3B/7B) failed at checkpoint load with a tensor
size mismatch (e.g. norm.scale 896 vs 1536 for 0.5B vs 1.5B). The LoRA
builder depends on the base model’s parameter size, not the family; the size
is now derived from the parsed base_id (Qwen2.5-0.5B → lora_qwen2_0_5b).
Validated on Apple-Silicon MPS end-to-end: the smoke LoRA now trains via the
local runner (5 steps, a real PEFT adapter — adapter_{0.pt,model.bin,config.json}
— written to outputs/adapter-<name>). Confirms the runtime venv + the pinned
torch/torchtune==0.4.0/torchao==0.5.0/kagglehub<0.3 set work on MPS.
0.1.0 — per-platform backend toolchain (BREAKING) + lineage aspect + de-shell
Breaking: lora_train no longer takes a backend = "..." attribute. The
training backend (local / runpod / modal) is now a per-platform toolchain —
select it with --platforms=@rules_lora//lora/backend:{local,runpod,modal}_platform
(default: runpod). lora_train resolves the toolchain for the jobspec composer +
backend identity; :<name>.run dispatches to the matching backend via select().
- Per-platform backend toolchain (
//lora/backend): abackendconstraint_setting + local/runpod/modal constraint_values, thelora_backend_toolchainrule +LoraBackendInfo, and convenience platforms parented on@platforms//host. - Lineage aspect:
lora_lineage_aspect(+LoraLineageInfo) threads training provenance up the dataset → recipe → base → train graph;lora_lineage(target=…)emits a JSON manifest — no hand-maintained sha plumbing. - De-shell: the local + merge
.runentries are nowpy_binaryorchestrators reading build-generated JSON configs;local_runner.sh+ the generated bash wrappers are gone. (Full hermetic torch vendoring + the orchestratorrunsubcommand are a follow-up — seedocs/hermetic-runners-roadmap.md.)
0.0.35 — Forward HF_TOKEN to the pod
- The synthesized manifest now always sets
forward_envs = ["HF_TOKEN"](plus"WANDB_API_KEY"when wandb is enabled). The pod’shf downloadof the base model needs the token for private or gated repos — e.g. a merged two-stage base (lora_merge→ private HF repo) or Llama. Before this, a private base failed setup with “repo is private, make sure you are authenticated.” runpod-cli skips anyforward_envsvar absent from the local env, so this is harmless when no token is set.
0.0.34 — lora_merge: fold an adapter into its base + export
- New
lora_mergemacro/rule. Folds a trained LoRA adapter into its base model (W' = W + (alpha/r)·B@A) and exports a standalone HF model dir —bazel run :<name>.run— with an optionalhfCLI push. CarriesLoraBaseModelInfo(id = push_repo)so the merged model is usable directly as alora_train(base = ...)(the two-stage rebase pattern: train a fluency adapter, merge it in, then train the task adapter on the fluent base). runtime/lora_merge(Rust, candle, CPU). The merge math, ported from a proven inference-time merge; validated byte-exact against peft’smerge_and_unload. Readsnum_hidden_layersfrom the baseconfig.jsonandr/lora_alpha/target_modulesfrom the adapter config; copies the base’s tokenizer aux files (vocab.json,merges.txt, …) so the result loads as a plain causal-LM. Addscandle-core+hf-hub(CPU-only, no GPU feature — the merge is I/O bound).- Currently merges attention projections (
q/k/v/o_proj); MLP modules and non-Qwen2/Llama key layouts are follow-ups.
0.0.33 — Fix single-file violation in synth manifest
lora_runpod_manifest_synthreturned two files inDefaultInfo(depset([out, dataset_jsonl])), which violated therunpod_manifestsrcallow_single_filecontract — every volume-path.runtarget failed at analysis with “must produce a single file.” The dataset only needs to be an action input (so it builds to bazel-bin for staging); it must not be inDefaultInfo. Now returnsdepset([out]).
0.0.32 — Network-volume data path for runpod training
lora_traingainsdata_volume/data_center. When set, the synthesized manifest mounts the RunPod network volume, stages the validated dataset to it via S3, reads the dataset from the mount (/workspace/...), tars the adapter into a single key, and drops the heavytraining/fullcorpus from the (now-skipped) workdir rsync. Pairs with rules_runpod 0.0.8’s volumestage/output_archive.- Fixes the genrule-dataset gap: datasets built by a genrule live only
in
bazel-bin, which the workdir rsync’sbazel-*exclude dropped — so they never reached the pod. S3 staging delivers them. Source-tree datasets are unaffected (legacy workdir path whendata_volume="").
0.0.31 — Fix gpu_type list arg passing
0.0.30 rendered the GPU fallback list correctly in the manifest but the
synth rule passed the candidates to the orchestrator as
--gpu-type A B C (flag once + bare values), which the orchestrator’s
clap Vec<String> rejects. Use add_all(..., before_each = "--gpu-type")
so the flag repeats per value (--gpu-type A --gpu-type B). 0.0.30 is
broken for multi-GPU runpod_gpu; use 0.0.31.
0.0.30 — GPU fallback list (rules_runpod 0.0.7)
lora_train’s runpod_gpu now accepts either a single GPU type
(string, unchanged) or an ordered fallback list, e.g.
runpod_gpu = ["NVIDIA L40S", "NVIDIA A40", "NVIDIA A100 80GB PCIe"].
The list is rendered into the synthesized manifest as
gpu_type = [...], and rules_runpod 0.0.7’s train tries each in turn,
advancing past capacity errors. Decouples the (frequently-changing,
availability-driven) GPU target from the model recipe — no more
hand-editing the BUILD and relaunching when SECURE capacity is dry. The
manifest synth emits the candidates via repeated --gpu-type; the
_lora_runpod_manifest_synth gpu_type attr is now a string_list.
0.0.29 — Detached training (rules_runpod 0.0.6)
Synthesized runpod manifest now sets detached = true +
poll_secs = 30, so the long-pole run script (the actual
tune run) executes detached on the pod and is polled for
completion rather than streamed over a tethered SSH session.
Three consecutive ~1hr agora parser fine-tunes died to mid-run
SSH Connection reset by peer. Training was converging each
time; the SSH session was the failure point. rules_runpod 0.0.6
adds the detached-execution primitives; this release opts the
training manifest into them. Bumps the rules_runpod dep 0.0.5 →
0.0.6.
0.0.28 — Wandb integration (runpod backend)
Opt-in W&B tracking on lora_train(backend="runpod"). Pattern
mirrors prime-transformer’s runpod-side wandb wiring:
lora_train(
name = "parser_full_jobspec",
...
backend = "runpod",
wandb_project = "agora", # NEW
)
When wandb_project is non-empty, the synthesized manifest:
- Sets
forward_envs = ["WANDB_API_KEY"]so runpod-cli propagates the local secret to the pod. - Adds
wandbto the pip-install in setup. - Runs
wandb login --relogin "$WANDB_API_KEY"(silent on success, warns + continues without W&B on failure or missing key). - Renders torchtune’s
metric_loggerasWandBLoggerwithproject = <wandb_project>andname = <adapter_name>.
Empty wandb_project (the default) is unchanged from 0.0.27:
no wandb pip install, no env forward, StdoutLogger only.
Caller needs to: have WANDB_API_KEY in the local env when
invoking bazel run :<job>_runpod_job.run. The key is forwarded
via SSH env, not baked into the manifest TOML.
0.0.27 — Align runpod torchtune pin with local + dump config
0.0.26’s runpod-side install pinned torchtune==0.3.1; local backend
uses 0.4.0. Same chat_dataset config that successfully trained
the 8-row seed on local-MPS produced zero iterations against the
85k-row corpus on the pod — the version skew turned out to be the
likely culprit. Bumped to 0.4.0 to match.
Also: the run script now cats the rendered torchtune YAML to
stderr and prints the first 200 bytes of the dataset’s first row
before invoking tune run. Helps diagnose dataset/config issues
when the Claude Code background-task buffer truncates the early
upload + setup output.
0.0.26 — Fix TOML escape in diagnostic line ($(pwd))
The 0.0.25 diagnostic line for the missing-dataset case used
\$(pwd) to defer shell-expansion, but TOML’s triple-quoted
basic string rejects \$ as an invalid escape sequence — the
manifest fails to parse and runpod-cli aborts before reaching
the pod. Drop the backslash: $(pwd) is fine, TOML passes it
through verbatim and bash expands it on the pod at run time.
0.0.25 — Re-publish of 0.0.24 (GitHub tarball cache churn)
The 0.0.24 tag landed correctly in git but GitHub’s archive endpoint returned 404 due to force-push tag-rewrite cache state. 0.0.25 is the same fixes, fresh tag.
0.0.24 — Runpod: explicit dataset path + correct outputs prefix (skip)
Two coupled fixes to the runpod backend that were causing the manifest synth to silently train zero iterations and drop the adapter on the pod:
-
Dataset discovery: the v0.0.23 run-script template walked
find . -name '*.jsonl'and picked the first file whose first 12 bytes contained"messages". On any workspace with multiple .jsonl files it would silently pick the wrong one, OR (when the actual SFT JSONL was filtered out of the rsync upload) leave DATASET empty and torchtune ran for zero batches. v0.0.24 bakes the explicit source path into the run script at build time, derived from the underlyinglora_dataset’ssource_path(new onLoraDatasetInfo). Errors loudly withpwd && ls -lacontext if the upload missed the file. -
Outputs prefix: the synthesized TOML had
outputs = ["adapter-<name>"]while the run script writes tooutputs/adapter-<name>/. The post-train rsync pulled the wrong path and silently dropped the adapter. Fixed tooutputs = ["outputs/adapter-<name>"]. -
New Starlark wiring:
LoraDatasetInfo.source_pathcarries the workspace-relative path of the lora_dataset’ssrc;_lora_runpod_manifest_synthreads it from the newdatasetattr (which providers-checksLoraDatasetInfo).lora_corpusconstructssource_path = ""since it’s a derived target with no single source JSONL — using alora_corpusdirectly in arunpod-backendlora_trainwill fail with a clear error.
0.0.23 — Pin torchtune to 0.4.0 (torch.cpu.memory_stats fix)
torchtune 0.5.0’s get_memory_stats calls
torch.cpu.memory_stats(), which doesn’t exist (only torch.cuda
torch.mpshave it). The call is unconditional even afterlog_peak_memory_stats=False. torchtune 0.4.0 (with torchao 0.5.0) doesn’t hit this code path and runs cleanly on Apple Silicon MPS.
Verified end-to-end: tune run starts on the agora parser smoke
without import-time or device-detection errors.
0.0.22 — Pin local-backend deps to a known-good triangle
The torchtune / torchao / kagglehub / kagglesdk dep graph breaks under several unpinned permutations on Apple-Silicon-MPS:
- Latest torchao (0.13+) requires
torch>=2.11. - Latest torchtune imports
from kagglesdk.kaggle_env import get_web_endpoint, which the 0.1.x kagglesdk on PyPI doesn’t export. - Latest kagglehub depends on a kagglesdk that breaks the import.
- torchtune 0.3.x doesn’t yet have the import path issue but
pre-dates
lora_qwen2_1_5bwe use.
Pin the local install to the May-2026 known-good set:
torchao==0.7.0torchtune==0.5.0kagglehub<0.3torch— let pip pick the latest matching version.
The RunPod backend continues to pin torchao==0.5.0 +
torchtune==0.3.1 in its own setup (matched to torch 2.4 in the
runpod/pytorch image).
0.0.21 — Local backend prefers python 3.11
torchtune’s transitive deps (kagglehub → kagglesdk) hit import-time
breakage under Python 3.14 (cannot import name 'get_web_endpoint').
Pick python3.11 if available (the ML stack’s lingua franca),
falling back to python3.12 then python3. On macOS with
brew install python@3.11 this picks the brew interpreter
automatically.
0.0.20 — Local backend: install torch + unpin versions
Two local-backend fixes uncovered by the agora smoke run:
-
Install
torchexplicitly in the venv (the previous pip install list assumed torch was already present, as in the RunPod pytorch image). Without it the MPS detection that runsimport torchalways falls through to CPU. -
Unpin
torchaoandtorchtuneversions for the local install. v0.0.17’s pin totorchao==0.5.0is not published for aarch64-apple-darwin (pipreturns 0.7.0 as the floor). Let pip pick the latest matching set per platform; the RunPod image still pins to torch 2.4-compatible versions in its ownsetup. -
Move device detection after the pip install so
import torchsucceeds and MPS gets picked on Apple Silicon.
0.0.19 — exec bash $RUNNER (no +x needed)
exports_files doesn’t stamp the executable bit on shell scripts;
v0.0.17/18’s generated wrapper did exec "$RUNNER" which failed
with Permission denied. Switch to exec bash "$RUNNER" —
bash reads the shebang directly without needing the exec bit.
0.0.18 — Fix _runfiles_path for external-repo short_paths
Cleanup-release of 0.0.17. _runfiles_path(file, ctx) returned
<workspace>/<short_path> unconditionally; for external-repo files
whose short_path already starts with ../<canonical>/..., this
produced runfiles paths like rules_lora+/../rules_lora+/runtime/...
that rlocation couldn’t resolve. Strip the leading ../ for
external-short-path inputs.
0.0.17 — lora_train(backend = "local") real entrypoint
The previously-stubbed local backend now emits a runnable
<name>.run sh_binary that drives torchtune on the host:
runtime/local_runner/local_runner.sh— venv bootstrap (workspace- local.venvs/lora-local/),pip install torchao==0.5.0 torchtune ==0.3.1 + HF tooling, HF base-model fetch, inline torchtune config render with auto-detecteddevice(mpson macOS,cudaifnvidia-smi, elsecpu),tune run lora_finetune_single_device --config <yaml>._lora_local_runnerprivate rule reads LoraRecipeInfo + LoraBaseModelInfo + LoraDatasetInfo and generates an entry script that passes the hyperparams positionally to the runner.lora_train(backend = "local")wires the rule into a<name>.runsh_binary.
Adapter lands at outputs/adapter-<name>/ in the user’s workspace
— no rsync, no pod, no orphan A100s. For 9-row LoRA fine-tunes
the on-laptop MPS path is the right default; reserve the runpod
backend for serious data volume.
0.0.16 — Manifest outputs path matches torchtune save dir
v0.0.15 declared the manifest’s outputs = ["adapter-<name>"]
but the synthesized run script saves to
$(pwd)/outputs/adapter-<name>. The mismatch silently failed
runpod-cli’s post-train rsync-back: the adapter trained, the pod
terminated (ephemeral), but the local outputs/ ended up empty.
Fix: outputs becomes ["outputs/adapter-<name>"].
0.0.15 — Drop o_proj default + ephemeral pod
Two paper-iteration QoL fixes:
-
lora_recipe.target_modulesdefault loseso_proj. torchtune’stune_to_peft_adapter_configdoesn’t accepto_projas a target module for Qwen2 / Llama3, so post-train peft conversion errored (with the adapter weights themselves saved fine). New default["q_proj", "k_proj", "v_proj"]round-trips through the peft config save. -
lora_train(backend = "runpod")now passesephemeral = Trueto the emittedrunpod_job. The.runwrapper threads--down-on-success --down-on-failurethrough runpod-cli, so an erroredtune runno longer leaves an orphan A100 burning $1.20/hr until manually deleted. Bumps min rules_runpod dep to 0.0.5 (where--down-on-failurelands).
0.0.14 — Drop unsupported model arg
lora_qwen2_1_5b doesn’t accept apply_lora_to_output. Remove
it; add lora_dropout: 0.0 for parity with the builder’s default.
0.0.13 — Use size-specific torchtune model builder
torchtune.models.qwen2.lora_qwen2 requires every architectural arg
(vocab_size / num_heads / embed_dim / …) explicitly. Switch the
qwen2 family to torchtune.models.qwen2.lora_qwen2_1_5b, the
torchtune-shipped builder that bakes the Qwen2/2.5-1.5B
architecture in. Hardcoded to 1.5B for v0; v0.0.14 generalizes via
a --family-variant flag.
0.0.12 — Round out the torchtune YAML config
torchtune 0.3.1’s lora_finetune_single_device recipe rejected
the v0.0.11 config with Missing key max_steps_per_epoch. Add
all of the keys torchtune unconditionally reads at recipe init:
max_steps_per_epoch: nullresume_from_checkpoint: Falsesave_adapter_weights_only: Trueenable_activation_checkpointing: Falselr_schedulerblock (cosine warmup, 1 step warmup)optimizer.fused: Trueprofilerblock (disabled)
0.0.11 — Pin torchao + torchtune versions
Latest torchao (0.13+) imports torch’s int1 dtype, which doesn’t
exist in torch 2.4 (the version in runpod/pytorch:2.4.0). Pin to
torchao==0.5.0 + torchtune==0.3.1 — last release where both
play nice with torch 2.4. Bumping the runpod image is a v0.0.12
follow-up.
0.0.10 — Add torchao to pod-side setup
torchtune now imports torchao unconditionally on package import.
Without it tune run exits with ModuleNotFoundError: No module named 'torchao'. Adds torchao to the setup’s pip install.
0.0.9 — Fix format string in v0.0.8 (yanked)
v0.0.8 shipped a Rust format! template with an unescaped {
inside a comment of the bash heredoc; the orchestrator binary
failed to compile. 0.0.9 swaps the offending JSON-snippet in the
comment for a prose description.
0.0.8 — Pod-side dataset auto-detect by content sniff (yanked)
v0.0.4–v0.0.7’s pod-side run block tried to find the SFT JSONL via
naming heuristic (*lora_dataset* or dataset.jsonl). That missed
common conventions like training/sft.jsonl and meant consumers
either renamed their seed file or wired in a bazel-bin path the
workdir rsync doesn’t see.
The detector now walks every .jsonl in the workdir (skipping
bazel-* dirs) and picks the first whose head bytes contain
"messages" — i.e. a real messages_v1 row. Robust to whatever
the consumer named the file. Failure message also names the
matched candidate count so debugging is one ssh away.
0.0.7 — Default RunPod image tag fix
v0.0.4 pinned runpod/pytorch:2.5.1-py3.11-cuda12.4.1-devel-ubuntu22.04
as the default RunPod image — but that tag was never published to
Docker Hub, so RunPod’s container daemon failed pod startup with
manifest unknown. Bump default to 2.4.0-py3.11-cuda12.4.1-devel- ubuntu22.04 — same Python/CUDA/Ubuntu stack, but a real tag (used
by prime-transformer’s working manifests).
0.0.6 — lora_train(runpod_cloud = "SECURE") knob
Adds a runpod_cloud attr to the lora_train macro (default
"SECURE"). Threads through to the synthesized manifest’s
[resources].cloud_type. SECURE is the right default for
paper-iteration runs — COMMUNITY tier is frequently exhausted
for popular GPU types (H100 / A100 / A40) and the resulting
create_pod: HTTP 500: There are no instances currently available
error is a poor first-run experience. Override to "COMMUNITY"
when cost matters more than availability.
0.0.5 — runpod manifest TOML structure fix
v0.0.4 emitted a TOML where setup and run followed [resources],
so the TOML parser folded them inside that table and runpod-cli
rejected the manifest with missing field setup. Fix: top-level
keys (name, workdir, outputs, setup, run) now come before the
[resources] table.
0.0.4 — lora_train pod-side manifest invokes real torchtune
The v0.0.2 write_file placeholder (run = """echo placeholder""")
is replaced with lora_runpod_manifest_synth, a private rule that
calls the Rust binary //runtime/runpod_orchestrator write-runpod-manifest. The Rust binary reads LoraRecipeInfo +
LoraBaseModelInfo and renders a manifest TOML whose setup and
run blocks:
- Install
torchtune,huggingface_hub[cli],transformers,datasetson the pod’s pre-baked pytorch image. - Pre-fetch the base model (revision-pinned) and stash the cached path for the train step.
- Render an effective torchtune config YAML inline by interpolating the LoRA hyperparams + per-job paths.
- Invoke
tune run lora_finetune_single_device --config ...— the real torchtune LoRA fine-tune loop. - Drop the adapter at
outputs/adapter-<name>/.
Supported model families today (selected by a family attr on the
synth rule, default qwen2):
- Qwen2.5 family (
qwen2) - Llama 3 family (
llama3) - Mistral family (
mistral)
Each maps to the matching torchtune tokenizer + lora model component. Extending the matrix is a single match arm in the Rust binary plus the corresponding tokenizer convention.
LoraRecipeInfo gains three new fields propagated from the
lora_recipe rule attrs:
learning_rate: str(kept as string so2e-4survives)micro_batch_size: intgrad_accum_steps: int
These were previously rendered only into the recipe YAML; the manifest synth needs them as structured attrs.
Smoke at examples/smoke/:
bazel build //examples/smoke:smoke_jobspec_runpod_manifest_toml
# renders a TOML whose `run` block has real `tune run` invocation
# instead of the v0.0.2 placeholder.
0.0.3 — lora_corpus rule with corpus-DAG deps
New public macro lora_corpus: declare an SFT dataset produced by
running a user-supplied transform binary over a source filegroup,
chained from upstream corpora via deps. Three consumers
(rules_agentic_ide chat traces, agora capability-auction corpus,
the NDA’d third) share the input/transform/validate/output skeleton;
the rule factors it out.
Public surface added:
lora_corpus(name, source, transform, deps, schema, min_examples)— runs the transform once with repeated--input/--corpus-dep/ single--outputflags, then runs the existingvalidate_jsonlvalidator on the transform output. ReturnsLoraDatasetInfo, so it plugs in anywherelora_datasetis accepted (in particular as thedatasetattr oflora_train).
Corpus deps form a DAG. The rule itself flattens to direct deps when invoking the transform (the upstream corpora’s transforms have already produced their validated JSONL artifacts as build outputs); Bazel’s build graph enforces no cycles and propagates transitive rebuilds.
Smoke at examples/corpus_smoke/:
bazel build //examples/corpus_smoke:derived_corpus
# base_corpus: 3 examples
# derived_corpus: 6 examples (3 source + 3 from dep)
Also includes the v0.0.2 features (deferred from registry release):
lora_train macro composes with @rules_runpod when
backend = "runpod", auto-emitting <name>_runpod_job.run.
0.0.1 — scaffold + public API frozen
Public surface (@rules_lora//lora:defs.bzl):
lora_dataset— typed SFT-JSONL dataset, validated + sha-pinned at build time. Schemas:messages_v1(OpenAI chat format),instruction_v1.lora_recipe— declarative training recipe. Frameworks:torchtune(rendered),axolotl(rendered),peft+trl(TODO templates).lora_base_model— HF hub model pinned by repo + revision.lora_train— composes the inputs into alora.v1.TrainingJobSpec(JSON, build-time) that a backend executes atbazel runtime.expert_manifest— bundles N adapters into theagentic_ide.v1.ExpertManifest.binpbshape (placeholder filegroup in v0.0.1; real rule in v0.0.2).
Runtime stubs:
runtime/torchtune_runner/{validate_jsonl, render_recipe}.py— build-time tools, std-lib only.runtime/runpod_orchestrator/(Rust) —write-jobspecsubcommand functional;runsubcommand pending v0.1.
Smoke at examples/smoke/ exercises all four macros end-to-end
and produces a self-contained jobspec — bazel build //examples/smoke:smoke_jobspec is the regression test.