Startrise Labs · Study
Study in progress · no results

Startrise Local Code: small coding models on a 32GB Mac

Startrise Local Code is our sister study to the frontier benchmark: it asks which small local coding models do the most useful work per second per GB on an M-series Mac, within a 32GB envelope. This page is the method. Results are pending.

One question: useful work per second per GB

tests passed ÷ wall-time (s) ÷ model size (GB)

Which small local coding models give the most useful work on Apple Silicon, with weights on an external SSD, inside a publishable memory envelope? The primary score is unit tests passing after a single-shot repair.

Every scored config has to fit a 32GB Mac

32 GBpublishable unified-memory envelope
≤22 GBpeak RSS hard limit (warn at 18)
pass@1primary metric · pass@3 secondary

Runs may be collected on a larger machine, but the claim we publish is the 32GB one. The pilot roster spans roughly 1.5B to 8B parameters, with a 14B model as an optional ceiling only.

Draft protocol

  1. Weights on the external SSD

    Verify the models path before any run.

  2. Repair fixtures

    Pilot: 2 Python + 2 TypeScript broken files with tests; more fixtures to follow.

  3. Single-shot repair

    Broken source + test names in, repaired file out, temperature 0.

  4. Run the tests, record telemetry

    Time to first token, tokens/s, peak RSS, wall time, size on disk.

  5. Aggregate

    Useful work / s / GB, with SSD vs warm-cache effects noted.

Pending

Status: pilot stage, no results

The harness, repair fixtures and pinned model digests are in place. No scored results exist in the study repo yet. This page will carry results when they exist, and not before.

Sources