Download README.md from DaisyChainAI/DaisyChain-Train: direct link, hf CLI and curl.
- Browser
- Download file 29.6 kB
-
https://huggingface.co/DaisyChainAI/DaisyChain-Train/resolve/main/README.md
- Command line
-
hf download hf://DaisyChainAI/DaisyChain-Train/README.md
-
curl -L -o README.md https://huggingface.co/DaisyChainAI/DaisyChain-Train/resolve/main/README.md
license: mit
tags:
- distributed-training
- old-hardware
- int8
- webgpu
- webrtc
DaisyChain-Train - old hardware training pipeline
Part of DaisyChain on Hugging Face: https://huggingface.co/DaisyChainAI
Model page (weights + card): https://huggingface.co/DaisyChainAI/DaisyChain-Train
I built this so spare / old machines can train a shared model. The training runs through emulated GPU logic: verified INT8 units (GUDA-style) that stand in for a GPU's math. Machines without a modern GPU still do the work. Chain several together and they train one model as a cluster.
Before you rely on it, read what it cannot do: Limitations.
Use the hardware you already have. Each machine runs the emulated GPU logic (verified INT8 units: multiply / requantize / ReLU). DaisyChain pools the machines data-parallel: device selection, capacity-weighted sharding, gradient sync, a P2P setup, and a live dashboard. Two ways to run: Docker or Python.
Read this first
DaisyChain-Train is for small models on spare hardware. It pools compute, not memory (the model must fit on one machine). Scaling is sublinear. It is not a substitute for a real GPU on large models. Full envelope: docs/LIMITS.md. Read that before you depend on it.
Feature list
Python cluster trainer (daisychain/)
- Data-parallel training across mixed machines. Each node trains its own shard. Gradients combine into the exact full-batch gradient. Replicas stay bit-identical.
- Capacity-weighted sharding. Faster machines automatically take a bigger share of the batch.
- Emulated GPU compute (verified INT8 units).
VerifiedLinearlayers run every forward multiply / requantize / ReLU through the bundled trained units. Rank 0 prints cluster-wide unit-invocation counts. - Bring your own model. Any
Task(build_model/sample/loss) viaDAISY_TASK. Template inexamples/my_task_template.py. - Plain-float alternative task. Same cluster and pooling with ordinary float math.
- Live dashboard (
daisychain-dashboard). Readiness banner, P2P connectivity scan, pooled cores/RAM, per-node capacity plan, live loss. - SpikeWhale control panel (
spikewhale_panel,localhost:8899). Sliders for model size / training settings, any HF dataset you can access (default streamed FineWeb-Edu), start/stop/re-adjust, live loss. - Docker demo cluster. 3 nodes + dashboard in one command.
- Windows helper (
scripts\setup.bat) and Tailscale mesh guide.
DaisyChain-Web (web/) - browser P2P training
- Zero-install nodes. Opening the page IS joining. Devices on one network auto-group (Snapdrop-style, by public IP).
- Private cross-network rooms.
?room=CODEwith host approval for every join. - Full WebRTC mesh. Gradients travel peer-to-peer. The server only signals and serves static files. It never sees weights or gradients.
- Leader-follower runs. Whoever presses Start sets width / sequence / batch-per-device / steps / learning rate for the whole group. Config is broadcast on the wire.
- Mid-run join. Late devices are synced in (weights + step) and contribute from the next step.
- Bit-identical replicas. Same seeded init, strict roster-order gradient averaging, deterministic Adam with identical state on every peer. Verified live by per-step weight hashes.
- Sync guard. Any weight-hash mismatch stops the run instead of training past a fork. The step roster forbids silent partial averages.
- Gradient repair. A follower missing a roster gradient re-requests it from the leader (8 steps retained), bit-exact, and the run continues.
- Cross-device kernel probe. Every step, every device re-hashes a fixed seeded int8 GEMM through its live kernel. Catches broken arithmetic that weight hashes cannot see.
- Hardcoded FineWeb-Edu streaming. The server reads random slices of the 10BT parquet shards straight off the HF CDN via HTTP range requests (pure-JS
hyparquet). Built-in corpus fallback offline. - Checkpoints. Download
.pt. Upload -> broadcast to the whole group. Validated (magic, dims, tokenizer vocab) before accepting. - Inference kit. One self-contained HTML file with the trained weights baked in. Generations offline, anywhere.
- In-page generation. Prompt box on the trained model.
- Old-hardware tier. No WebGPU? The identical units run on CPU (same bits, so CPU and GPU devices co-train in one group). There is no plain-float path.
- Large-message fragmentation. Multi-MB gradients/checkpoints chunked at 48 KB over the data channels.
Verified compute and kernels (web)
- Verified INT8 units everywhere. Block-scaled int8 GEMM: exact LUT products, exact int32 accumulation, bit-exact f32 epilogue with a pinned rounding schedule. Scales derived in JS f64 (division never runs on GPU).
- Backends, best-first. DP4A hardware int8 dot -> LUT compute shader -> CPU mirror. Every kernel is exact-gated at init (bit-level compares) and demoted to the mirror on any mismatch.
- Continuous random-cell audit at live training shapes.
- Fused attention kernels. Gather/scatter head-strided q·kᵀ and a·v straight from BT×C layout (CUTLASS ex. 36/52 style).
- QKV dual-GEMM fusion. Shared left operand quantized once, one batch-3 dispatch (ex. 45). Bit-identical.
- B2B MLP chain. Both MLP GEMMs back-to-back on GPU with fused per-row absmax reduction and on-device quantize (ex. 13 + 23). WGSL-exact respec with a fround-stepped JS mirror. FMA-contraction-immune by construction.
- Dispatch-optimized backward. Overlapped independent GEMMs, batch-3 sibling fusions (ex. 05/24). Bit-identical gradients. Optional int8 STE backward path (dormant, 1.21x vs float).
Verification stack (web)
- Exact init gates on every kernel, every device, every boot. Including gates that "gate the gate" with discriminating boundary inputs.
- IEEE-754 binary32 oracle in exact BigInt arithmetic. Proves the JS epilogue mirror is spec-correct (rejects the old mirror on 34% of inputs).
- Metamorphic property suite. Reference-free relations + definitional absolutes. 4/4 on an externally-authored bug corpus, matching the differential gate.
- RDNA2 ISA audit hardenings. Bit-level (-0-aware) gate comparisons. Proof that FMA contraction cannot change the quantize.
- Twelve-suite test chain (
cd web && npm test): convergence, replicas, oracle, gates, properties, external corpus, self-corpus (the instruments scored against my own bugs), B2B, optimizer, transformer LM, int8 backward, non-finite quantize refusal. Results in web/TEST_RESULTS.md. - Dirty-buffer gate. The pool is poisoned before a re-sweep so state bugs (a kernel assuming zeroed memory) are caught deterministically rather than by ordering luck.
Documentation
- Python: QUICKSTART, LIMITS, CUSTOM_TASK, TAILSCALE.
- Web: Getting started, Architecture, Verification, Troubleshooting.
Quick start
Docker (most reliable - one command)
docker compose -f docker/docker-compose.yml up --build
# open http://localhost:8080
Brings up a 3-node demo cluster + dashboard on one machine.
Python (real machines)
On every machine (pip install -e .):
export MASTER_ADDR=100.101.102.10 # coordinator IP (Tailscale 100.x recommended)
export MASTER_PORT=29560
export WORLD_SIZE=3
export RANK=0 # 1, 2, ... on the others
export GLOO_SOCKET_IFNAME=tailscale0 # your mesh / LAN NIC
daisychain-train
Windows helper
scripts\setup.bat
An interactive menu: Docker, Python node, or just install deps.
Full walkthrough: docs/QUICKSTART.md.
SpikeWhale control panel (sliders -> real training)
python -m daisychain.spikewhale_panel
# open http://localhost:8899
A web control panel: pick model size / training settings with sliders, choose any Hugging Face dataset you have access to (default: streamed FineWeb-Edu), hit Start, and watch the live loss. Stop and re-adjust any time with <- Back to settings. Launches the real DaisyChain training underneath.
DaisyChain-Web (train by opening a browser tab)
cd web && npm install && node server.js
# open http://localhost:8787 on every device
Zero-install browser training. Devices on the same network auto-group (Snapdrop-style) and train a shared model peer-to-peer over WebRTC, computing through the same verified INT8 units (WebGPU, with the identical units on CPU for machines without it - there is no plain-float path). Private cross-network rooms via ?room=CODE with host approval. The room creator accepts each device before it can join. Includes gradient averaging with a deterministic Adam optimizer (identical state on every peer, nothing extra over the wire), checkpoint download (.pt) and upload -> broadcast so one device can restore the whole group after a failure.
Live demo: https://huggingface.co/spaces/Quazim0t0/DaisyChain-Web
How it works
Each machine runs the same command. They form a cluster and train one shared model. Two things happen:
- The compute runs through the emulated GPU logic. By default the model is built from
VerifiedLinearlayers, so every forward multiply / requantize / ReLU is done by the bundled verified INT8 units (daisychain/verified/) - the emulated GPU math. Rank 0 prints cluster-wide unit-invocation counts so you can see the emulated logic doing the work. - The machines are pooled data-parallel. Each node trains on its own shard. Gradients are capacity-weighted and combined into the exact full-batch gradient, so replicas stay bit-identical. Faster machines automatically take a bigger share.
old machine A -+
old machine B -+- each runs the emulated GPU logic on its shard -> one model
old machine C -+ (gradients combined across the cluster)
Bring your own model
DaisyChain-Train trains any Task (build_model / sample / loss). Copy examples/my_task_template.py, set DAISY_TASK=your_module:YourTask. Use VerifiedLinear (see daisychain/verified_task.py) to run your model's compute through the emulated units. See docs/CUSTOM_TASK.md.
Plain-float alternative
To skip the emulated units and train with normal float math on each machine, set DAISY_TASK=daisychain.example_task:ExampleTask. Same cluster, same pooling. The model math just runs as ordinary float instead of through the verified units.
The dashboard
daisychain-dashboard (or the Docker service) serves a Tailwind page at :8080: readiness banner, P2P connectivity scan, pooled cores/RAM + capacity plan (per-node device, weight, batch), and live training loss.
Networking
Use Tailscale for a P2P mesh so machines on different networks get stable IPs on one interface. docs/TAILSCALE.md.
Layout
daisychain/cluster.py capacity-weighted CPU/GPU data-parallel trainer
daisychain/train.py entry point (daisychain-train)
daisychain/verified/ bundled trained N/N units + VerifiedLinear (train through them)
daisychain/verified_task.py default task: forward runs on the verified units
daisychain/example_task.py plain-float alternative task
daisychain/task.py the Task interface + loader
daisychain/dashboard/ agent + P2P scanner + Tailwind server
docker/ Dockerfile, dashboard image, compose (demo cluster)
scripts/setup.bat / setup.sh interactive setup helpers
config/ nodes + cluster env examples
examples/my_task_template.py starting point for your own model
test_verified_units.py full-domain checks of every shipped verified path (25)
docs/ QUICKSTART, LIMITS, CUSTOM_TASK, TAILSCALE
daisychain/spikewhale_task.py trains the real SpikeWhale on streamed HF datasets
daisychain/spikewhale_panel.py slider control panel (localhost:8899)
web/ DaisyChain-Web: P2P browser training (WebRTC + WebGPU)
export_luts_web.py regenerates web/public LUTs from the trained units
Recent updates (30 September 2026) - reductions refuse what they used to accept
Six fixes, each found by applying a lesson from differential-testing an RDNA2 shader translator: a second implementation nobody checks is not verified, inactive or absent participants must not leak into a reduction, a max over NaN is undefined, and know your accumulator width. Every one is covered by a check that fails on the old code. python test_verified_units.py + python test_cluster_reductions.py and cd web && npm test (thirteen suites) reproduce all of it.
| # | where | what was wrong | now |
|---|---|---|---|
| 1 | browser, WebGPU MLP chain | the row max comes from an atomicMax over bitcast(abs(h)) and never passed the non-finite guard the CPU path has; a NaN/Inf h1 became inv = 0/NaN and the GPU quantized floor(NaN) (implementation-defined) where a CPU peer threw |
scalesFromAbsMax refuses a non-finite max. NaN and Inf bits order above every finite value as u32, so the row max is non-finite iff some element is - one check covers the whole row (checked on 2000 random rows) |
| 2 | browser, averageGrads |
a gradient of a different length was silently averaged over a prefix (typed arrays ignore out-of-range writes); a NaN/Inf gradient went straight into DaisyAdam, poisoning m/v on every peer identically, so no fork check could see it |
refuses both with GradientError; the training loop halts at that step with weights and Adam state untouched (still a valid checkpoint) |
| 3 | browser, int8 GEMMs (JS and WGSL) | the Int32Array and WGSL i32 accumulators wrap past K = 131,071 - CPU and GPU agree, on a wrong sum |
refused with AccumulatorBoundError at every GEMM entry point; the same bound as the Python kernel's MAX_INT32_K |
| 4 | Python, logged cluster loss | sum(loss_i) / world - an equal-weight mean of unequal, capacity-weighted batches |
cluster_mean_loss = Σ wᵢ·lossᵢ with the gradient's own weights. Measured on 3- and 13-sample nodes: old 1.418, true full-batch loss 0.959, new 0.959 |
| 5 | Python, gradient all-reduce | one node's NaN was summed into every replica and applied | check_finite_grads refuses the reduced gradient before opt.step(); every rank holds the same sum, so every rank stops at the same step with no extra communication |
| 6 | Python, pick_device |
with CUDA_VISIBLE_DEVICES set but empty (a common "CPU only"), torch reports is_available() = True and device_count() = 0, and the node crashed at startup |
a device must actually exist; falls back to CPU |
Unchanged for valid inputs: finite gradients, finite activations, and K <= 131,071 produce bit-identical results to before, so updated and older peers still co-train. A 2-rank end-to-end DaisyCluster.fit run stays bit-identical across replicas (replica_diff 0.0). The WebGPU-side changes are host-side JavaScript (checked by node --check and by the Node suites through the shared mirror functions). The WGSL kernels themselves are unchanged.
Recent updates (September 2026) - requant rounding, bounded units, non-finite guards
Everything below is measured. python test_verified_units.py (25 checks) plus cd web && npm test (twelve suites) reproduce it.
The requant rounds instead of truncating. NeuralRequant16 was sat_int8(x >> 8), and Python's >> on negatives is an arithmetic floor, so every requantized activation carried a systematic bias. Over the full int16 domain:
| requant | mean error | disagrees with the other on |
|---|---|---|
floor x >> 8 (old) |
-0.4981 LSB | 32,640 / 65,536 |
round-half-up (x + 128) >> 8 (now) |
+0.0020 LSB |
The shipped requant16.pt implemented the floor exactly, so this was real behaviour, not a stale docstring. It was retrained against the new reference and saved only at 65,536/65,536 (a 99.9% unit would silently break the LUT equivalence the browser relies on), and web/public/requant_lut.bin was re-exported to match. The invariant is neural forward == LUT == native integer op, so a table that lags its unit is a fork. Tie rule, stated precisely: this is half-UP. Python/numpy/torch round is half-to-even. The two differ on exactly the 128 exact ties (0.2% of inputs) and both are unbiased to within 0.002 LSB. The Python fleet quantizes with np.round and the browser with floor(x + 0.5). The two fleets never co-train, so that cannot fork a group.
Bounded elementwise units. The GEMM backends already capped their temporaries; the units right after them did not. relu_array / requant_array received the whole activation matrix in the proof path at ~1.5-2 KB per element (262,144 elements = 407 MB), and the multiply-LUT build pushed 262,144 rows through the atom net at startup (277 MB RSS on every node). Both are now blocked: 407 MB -> 2 MB, flat with N; 277 MB -> 1 MB. Elementwise blocking is exact (no accumulation order), verified bit-identical at N = 1 / 999 / 32768 / 32769 / 100000. dataset() builders are vectorised (65,536 Python iterations -> one numpy pass, 0.015 s), which is what makes retraining a unit practical at all.
A NaN or Inf no longer quantizes to zeros. A float -> int8 conversion has no answer for NaN/Inf, and both trainers gave a silent wrong one:
| path | input | old result |
|---|---|---|
Python qat.py |
[0.5, nan, 3.0] |
scale = NaN -> [0, 0, 0] (only a RuntimeWarning) |
Python qat.py |
[0.5, inf, 3.0] |
scale = Inf -> [0, 0, 0] |
| browser quantizers | [0.5, Inf, 3] |
scale = Infinity -> [0, 0, 0] |
| browser quantizers | [0.5, NaN, 3] |
NaN skipped by the |max| scan, stored as 0 |
One bad value zeroed the whole layer's int8 input and training carried on. Now VerifiedLinear raises FloatingPointError, and every browser quantizer (quantize, quantizeRows, quantizeCols, rowAbsMax, quantizeHeadCols, and the transformer's quantizeColsAsRows) throws NonFiniteError, naming how many values were bad. Finite inputs are untouched, so builds with and without the guard still co-train. Both new tests fail against the old code (Python: 2 checks; web: test_nonfinite.js, 12) and pass now.
Recent updates (August 2026) - verified compute path
Three changes to daisychain/verified/, each measured rather than asserted. python test_verified_units.py (24 checks) reproduces all of it.
Bounded GEMM memory. LUTBackend.gemm and NeuralBackend.gemm materialized the whole (m, n, k) product block before reducing it, so peak memory was cubic in layer width. At the SpikeWhale panel's maximum (hidden 768, sequence 512) a single layer's forward allocated 2.4 GB - on the spare hardware this project exists to use. Both now block the m axis. The contraction axis k is untouched, so the sum and its order are unchanged and results stay bit-identical:
| GEMM | before | after |
|---|---|---|
| 64x64x64 | 2.3 MB | 2.3 MB (unchanged path) |
| 256x256x256 | 135.8 MB | 68.9 MB |
| 512x768x768 (panel max) | 2426.9 MB | 77.2 MB |
| neural backend, 64x96x96 | 519.2 MB | 57.5 MB |
Caps are class attributes (LUTBackend.max_block_bytes, NeuralBackend.max_products), tunable per instance. Below the cap the LUT path takes the original single-block branch unchanged.
The two backends need different accounting, which is worth knowing before tuning them: a LUT product costs 8 bytes, but a neural product is split into four nibble pairs through a 128-wide net and costs ~880 bytes. Applying the LUT's byte budget to the neural path computes a block 110x too large and never splits at all.
The multiply LUT is certified, not trusted. build_luts() returned its tables unchecked while the docs described a self-certify gate. Because the table is the complete finite domain, certifying it is the exhaustive verification: LUTBackend.__init__ now checks all 65536 entries against signed integer multiply and raises on mismatch. Cost: 0.5 ms.
Full-domain tests for the paths that ship. The units' exhaustive verification covers the scalar entry points (NeuralMul8.verify() walks mul()). Production runs the batched ones: mul_array, relu_array, requant_array, and the LUTs built from them, which are different code. test_verified_units.py checks every shipped path against golden integer arithmetic over its complete domain. All pass today. Nothing would have caught a regression before.
Two guards against silent no-ops. gemm_int8 enforces the int32 accumulator bound (MAX_INT32_K = 131071) instead of asserting it in a comment, since layer width is user-selectable and an overflow would produce wrong numbers rather than raise. instrument.require(**minimums) fails when a verified unit was under-invoked, and refuses to report at all while counting is disabled. A zero from a probe that was switched off is not evidence that the units did not run.
Recent updates (July 2026) - DaisyChain-Web
Verification stack. The browser trainer's correctness is now checked by things that run, not argued. Full results: web/TEST_RESULTS.md.
- IEEE-754 oracle (
web/test_ieee.js). A binary32 oracle built from the standard in exact BigInt arithmetic proves the JS epilogue mirror is spec-correct, and rejects the old round-once mirror on 34% of inputs. - Metamorphic properties + oracle mutation scoring (
test_metamorphic.js,test_corpus.js). Properties needing no reference implementation, scored against an externally-authored bug taxonomy. 4/4, matching the exact differential gate's 4/4. Relations own the loop bugs. Two definitional absolutes (ReLU output range, a unit-scale integer anchor) own the value bugs no relation can see. - Exact kernel gates on every live kernel, a continuous audit at live shapes, and a cross-device kernel probe (same seeded int8 GEMM, same hash on every honest device, any backend). The audit's sampling was rebuilt against a named bug class. Its old constants (6 cells, 2% of GEMMs) bounded an overhead that had never been measured (auditing every GEMM costs <0.01% of a step), and uniform random cells cannot see a last-row/column bug at a 16512-wide output. Sampling is now stratified: the first cells are the structural danger points, chosen deliberately. On a last-column bug, same cell budget: uniform caught 5/300 audits, stratified 300/300, with zero false positives.
- RDNA2 ISA audit. Reading a real GPU's shader ISA against our determinism assumptions confirmed three of them on silicon (exact packed int8 dot; correctly-rounded f32 add/mul; 1-ULP reciprocal - division stays off the GPU) and produced two hardenings. (1) Real ISAs have non-IEEE variants that flush -0 to +0. JS
!==cannot see that (-0 !== 0is false), so all gates and audits now compare bit patterns - exactly what the replica hash sees. (2) FMA contraction of the quantize'sx·inv + 0.5(one rounding instead of two) turned out to be floor-invisible by construction - proven intest_b2b.jswith 175k+ last-ulp anomalies at binade edges, zero survivingfloor(). Rounding mode and denorm flushing are runtime driver state on real hardware, which is why every device re-runs the exact gates at every init.
Training data. FineWeb-Edu (10BT sample) is the hardcoded dataset. The Space reads random slices of the parquet shards straight off the HF CDN with range requests (pure-JS hyparquet, SNAPPY) and serves plain text at /data. No dependency on the datasets-server rows API and its 503s.
Resilience. The sync guard now repairs instead of halting. A roster gradient that reached the leader but not some follower (asymmetric WebRTC mesh) is re-requested from the leader, bit-exact, and the run continues. The guard still stops anything that would fork the weights.
CUTLASS-style kernel work, each step proven bit-identical or exact-gated:
- Dispatch-optimized backward (ex. 05/24): independent GEMMs overlapped, sibling trios fused into batch-3 dispatches. Bit-identical gradients. Dormant int8-backward path down from 1.63x to 1.21x vs float.
- QKV dual-GEMM fusion (ex. 45): q/k/v share one left operand. Quantized once, one batched dispatch, zero changed bits.
- B2B MLP chain (ex. 13 + 23): both MLP GEMMs back-to-back on the GPU with a fused per-row absmax reduction. The intermediate is quantized on-device via a WGSL-exact respec (
floor(f32(x·invScale)+0.5)- no GPU division) whose fround-stepped JS mirror keeps mixed GPU/CPU fleets bit-identical.
Profile-driven speed work. Every change below is bit-identical (gradient and loss hashes unchanged), so none of it trades correctness for wall clock:
- Buffer pooling. GPU buffers are recycled by size bucket instead of being created and destroyed per dispatch (~19 per MLP call, per layer, per step). 6-10% faster, every hash unchanged.
- Shared-operand embedding GEMMs. Profiling put two f32 backward GEMMs at 55% of the entire step, and both consumed the same
dlogitsoperand: ~17 MB at the 16512-token vocab, uploaded twice. One upload, one encoder, one submit: that pair went 205 -> 90 ms and the step 12% faster. The fusion is gated bit-for-bit against the two calls it replaces. - A negative result, kept on purpose. The remaining hot kernel looked cache-hostile (adjacent lanes wrote 66 KB apart), but making the writes contiguous changed nothing. Two probes explain why: holding the output at 17 MB while cutting compute 32x barely moved the time. The logits GEMM is transfer-bound, not compute-bound, and the readback cannot be removed because softmax must stay in JS (WGSL's
expis not correctly rounded, and a per-vendorexpwould fork replicas). The real lever there is the vocabulary, not the kernel. - An init backend race was tried and removed. It tied on the shipped path, cost ~430 ms of init, and made the backend vary between page loads, which silently invalidated three A/B comparisons before it was caught. A knob that changes what you are measuring is worse than a fixed choice.
Dirty-buffer gate - and the assumption it falsified. Pooling introduced a bug class the gates predate: a pooled buffer is not zero-initialized, so a kernel that assumes zeros is right on step one and wrong on step two. That is a state bug, where no single call is wrong and the sequence is, which is the family no oracle can reach. The assumption was that the gates were blind to it. Mutation-testing the gate proved otherwise: deleting the zeroing made the plain gate fail at its second shape, because the sweep's own shapes recycle each other's buffers. The suite had incidental coverage nobody designed, which is coverage nobody can rely on. Shorten the shape list and it evaporates with the gate still green. It is now deliberate: the pool is poisoned with 1e4-magnitude residue before a re-sweep, so detection no longer depends on ordering luck. ~90 ms one-time.
Scoring the oracles against my OWN bugs (web/test_selfcorpus.js). The external corpus measures kernel bugs someone else wrote down, so this suite asks the harder question: what do the instruments score against the four real bugs of the month? Properties 0/2 on the data-plane pair (the c·out theorem again), differential 2/2 - but half the bugs were not in the kernels at all. A dead gate is a bug in a checker, caught only by mutating the gate. A stalled roster gradient is a bug in the protocol, where every computed value on every peer was correct, so no data oracle could fire. Those needed different instruments, not better oracles.
All thirteen test suites (cd web && npm test) pass. Results with methodology in web/TEST_RESULTS.md.
Install
pip install torch numpy psutil
pip install -e . # exposes: daisychain-train, daisychain-agent, daisychain-dashboard
Requires Python >= 3.9, PyTorch >= 2.0. Multi-node is reliable on Linux/macOS. On Windows use Docker/WSL (see Limitations).
Links
- DaisyChain on Hugging Face: https://huggingface.co/DaisyChainAI
- This model: https://huggingface.co/DaisyChainAI/DaisyChain-Train
License: MIT · Author: Dean Byrne (Quazim0t0) · Org: DaisyChainAI
Citation
@misc{byrne2026daisychain,
title = {DaisyChain-Train: An Old Hardware Training Pipeline},
author = {Byrne, Dean (Quazim0t0)},
year = {2026},
howpublished = {\url{https://huggingface.co/DaisyChainAI/DaisyChain-Train}},
note = {Chain spare/old machines into a data-parallel training cluster}
}
Dean Byrne (Quazim0t0) · 2026