Products · Pre-alpha
Ogun logo

Forge raw sources into
reproducible datasets.

A local-first developer platform for data mining, data engineering and agent benchmarking, built on the Zeus download engine.

Status
Pre-alpha
Stage detail
Contracts phase
Demo
On request
Built by Verne
Yes
Pre-alpha

Contracts phase. Architecture decided, specifications in progress, no code and no release yet.

Exists today

  • Two architecture decisions are accepted. Three more are proposed and described here as designed, not settled.
  • There is no code. The repository holds documents only.
  • The Ogun specifications are not written yet.

Planned, not built

  • Work on the runtime starts only after the Zeus engine passes crash-and-resume. It has not.
  • Later stages add tool adapters, a model ladder, a benchmark profile and a desktop studio. Targets, not promises, and none have dates.
  • Before any public launch the name needs trademark clearance and a cultural review.
  • Apache-2.0, source opens at alpha
  • Source is private
  • Not released
What it is for

The forge for reproducible data.

AI, LLM and data engineers who need to turn messy web and filesystem sources into verified, reproducible datasets on a schedule, and to let agents plan that work while people hold the approvals.

01

A long collection job dies and you start over.

Designed to resume verified acquisition after a crash or reboot, redoing only what is incomplete.

02

A dataset arrives and nobody can say how it was made.

Designed to store artifacts by content with a lineage record, including the model and prompt hash where one was used.

03

An agent writes a shell command and you run it.

Proposed route: prompt, then a workflow description, then a validator, then a plan diff, then a human approval. Never prompt straight to shell.

04

Jobs must run on a schedule, with gaps filled in later.

Designed with triggers, partitions and backfills as first-class ideas.

05

Tools drift between runs.

Designed to run tools such as DuckDB, dbt, ffmpeg and OCR as sandboxed operators with pinned versions.

See the idea

Eight stations, one approval in the middle.

A picture of the intended lifecycle. It is a design sketch, not a recording of a run, because there is no runtime.

Lifecycle · planned, concept
Climate Papers · run 184running

Step 1 / 9 Acquire: sources are fetched with resumable, verified transfers.

Planned, concept. The run name and the OCR question are an illustration of the approval step. Approvals are designed to be answerable only by a person.
How it works

8 steps, in plain words.

Steps describe the design. Where a step is not built yet, the status section above says so.

  1. 01

    Acquire

    Fetch sources with resumable, verified transfers.

  2. 02

    Materialize

    Turn what arrived into stored, content-addressed files.

  3. 03

    Verify

    Check hashes, types and expected shape before going on.

  4. 04

    Transform

    Run pinned tools to clean and reshape.

  5. 05

    Organize

    File the results into a layout you can query.

  6. 06

    Extract

    Pull structured fields from documents, with OCR when approved.

  7. 07

    Curate

    Select, label and deduplicate for the dataset you want.

  8. 08

    Publish

    Write the dataset out with its lineage attached.

Principles

What we hold it to.

  • Agents do the planning. Humans hold the approvals.
  • Deterministic infrastructure under probabilistic intelligence.
  • A workflow with no model step is designed to run with no model access at all.
  • Integrate existing tools, do not replace them (proposed).
  • Every byte is meant to have a lineage, and every run is meant to be repeatable.
  • No telemetry, no captcha circumvention, no reading browser cookies.
Demo available on request

See Ogun with the team that is building it.

Design walkthrough, no runnable build.

Pre-alpha. Not released. Source is private.