Local-first capture
Captured on your box and PII-redacted at source, fail-closed. Runs fully offline; your traces land as open ATOF JSONL you own and can export. --remote is opt-in.
train.cloud is the self-improving loop for coding agents: capture how your team actually codes, train an open model on your own verifiable signal — your CI, your tests — and serve it back, retrained continuously as you ship. You reach it through Border Collie, its free open-source front door: a single binary that captures locally, PII redacted on your box, and connects you to the loop with one command.
cargo install bcollie--remote is opt-in. Read the docs →A real layout, not an image — click a tab. Capture is the panel that is live in the binary today; the rest print exactly what you see here until they ship.
One open funnel from your real work to a tuned model — and one base-URL flip back onto your own tools. Switch isn't the finish line: you keep coding, and the traces you capture next feed the next round, so the model keeps improving as your codebase moves.
Border Collie has no runtime of its own — it drives your harness, never replaces it. Execution stays in klein or your CLI; Border Collie captures, mines, and coordinates, then hands the two commands that matter — train and deploy — to train.cloud across one API. A new model ships only if it beats the last. And the loop doesn't stop at switch — it keeps capturing and retraining, so the model that fits today keeps fitting as your code drifts.
Your agent's traffic, captured locally and PII-redacted on your box.
Recurring workflows become training tasks with verifiers — from your traces.
SFT + RL from verifiable rewards (your CI, your tests), eval-gated.
A served, OpenAI/Anthropic-compatible endpoint tuned on your work.
Flip one base-URL — same UX, better model. Then keep coding; new traces start the next round.
Local-first capture, driving your own harness, training on your own verifiable signal, with a one-flip switch back onto your tools. Any one exists somewhere; the combination doesn't.
Captured on your box and PII-redacted at source, fail-closed. Runs fully offline; your traces land as open ATOF JSONL you own and can export. --remote is opt-in.
Not just observes — drives klein, Claude Code, Codex, Cursor, or Gemini through one adapter. One experience across every harness and four surfaces: CLI · TUI · WebUI · desktop/mobile.
The training set comes from your work, with provenance and verifiers — not a generic corpus. SFT + RL from rewards you can grade (your CI, your tests), eval-gated before anything ships.
After deploy, repoint one base-URL and your harness runs the tuned model — same UX. Default is a served endpoint; export the weights you own on the enterprise tier.
Fine-tune once and the model that fit your codebase in March is stale by June — your code drifted, the model didn't.
train.cloud's loop retrains on your newest traces as your codebase drifts, amplifies your signal with self-play you can't reproduce offline, and eval-gates every version so it only ships if it beat the last. That's the engine behind the seam — the part a fork of the open front door can't clone.
Nothing leaves your machine unless you run train/deploy. Everything in the front door is Apache-2.0 and never relicensed — the capture proxy, importers, local task-mine, orchestration, all four surfaces, the offline mock, BYO-harness — and it all runs offline. Only the managed train/deploy loop behind the seam is paid.
Your private edge stays in your tenant; only de-identified, privacy-bounded signal is ever pooled — opt-out, and enterprise can disable it entirely.
Capture, mine, orchestrate, all four surfaces, offline mock. Real tooling, not a trial.
train/deploy: self-play, eval-gate, managed serving. Named at the boundary, never reimplemented.
Nothing leaves until you train/deploy — and even then, redacted at source. Federated is opt-out; enterprise can disable it.
Capture/eval clouds (LangSmith, Langfuse, Braintrust, Helicone) record your traffic and stop at a dataset. Tuning platforms (OpenPipe, Predibase, Fireworks RFT, Axolotl, Tinker) fine-tune — but only from a dataset you hand them. Neither drives your harness, and neither switches you back. Border Collie is the connective tissue.
| Capability | Border Collie | Capture / eval clouds | Tuning platforms |
|---|---|---|---|
| Local-first capture, redact-at-source, offline | Yes | Partial | No |
| Drives your real harness (not just observes) | Yes | No | No |
| Training set mined from your traces | Yes | Dataset export | Bring your own |
| Closes the loop: train → deploy a tuned model | Yes | No | Yes |
| One-flip switch back onto your tools | Yes | No | No |
| Apache-2.0, local-first, portable ATOF | Yes | Varies | Library-only |
Any one of these exists somewhere. The conjunction — one local-first, harness-agnostic, Apache-2.0 funnel — doesn't.
The whole open funnel runs offline with no account. train/deploy against train.cloud is in private beta with design partners.
For teams running an agent harness at real volume who want it cheaper, self-hostable, and tuned to their own code.
Quickstart, the training-loop internals, self-hosting, and the privacy model — no hand-waving.
Ships with the open-source release.
cargo install bcollie — capture your first workflow offline in minutes.
How your traces become a tuned model, and what crosses the one seam.
Deploy on your own GPUs, air-gapped if needed. Export the weights you own.
What's pooled, what never leaves your tenant, and the opt-out.
Drop-in wire dialect; usage, credits, and rate limits.
Drive Claude Code, Codex, Cursor, or Gemini through one adapter contract.