A long collection job dies and you start over.
Designed to resume verified acquisition after a crash or reboot, redoing only what is incomplete.
A local-first developer platform for data mining, data engineering and agent benchmarking, built on the Zeus download engine.
Contracts phase. Architecture decided, specifications in progress, no code and no release yet.
AI, LLM and data engineers who need to turn messy web and filesystem sources into verified, reproducible datasets on a schedule, and to let agents plan that work while people hold the approvals.
Designed to resume verified acquisition after a crash or reboot, redoing only what is incomplete.
Designed to store artifacts by content with a lineage record, including the model and prompt hash where one was used.
Proposed route: prompt, then a workflow description, then a validator, then a plan diff, then a human approval. Never prompt straight to shell.
Designed with triggers, partitions and backfills as first-class ideas.
Designed to run tools such as DuckDB, dbt, ffmpeg and OCR as sandboxed operators with pinned versions.
A picture of the intended lifecycle. It is a design sketch, not a recording of a run, because there is no runtime.
Step 1 / 9 Acquire: sources are fetched with resumable, verified transfers.
Steps describe the design. Where a step is not built yet, the status section above says so.
Fetch sources with resumable, verified transfers.
Turn what arrived into stored, content-addressed files.
Check hashes, types and expected shape before going on.
Run pinned tools to clean and reshape.
File the results into a layout you can query.
Pull structured fields from documents, with OCR when approved.
Select, label and deduplicate for the dataset you want.
Write the dataset out with its lineage attached.
Design walkthrough, no runnable build.
Pre-alpha. Not released. Source is private.