Compress

Compress a model

Your compressed checkpoint, measured on your evaluation — same architecture, fewer parameters. We send a proposal.

Offering

A supervised compression run

We schedule and supervise the run: one compression level on your trained checkpoint, on high-performance compute we provision with our hardware partners, using methods from our own ongoing research.

  • Compressed checkpoint — weights at the agreed density, same architecture as your original
  • Plan file — the connectivity configuration and run parameters used for the compression
  • Evaluation card and report — side-by-side results of original versus compressed on your named evaluation

Process

How it works

  1. Plan or share your model

    Use the planner on a Hugging Face model, or describe your checkpoint here.

  2. Request a proposal

    Tell us your data and evaluation. We review manually and send a proposal.

  3. We run the compression

    After you accept, we schedule the compute and supervise the run — with regular progress reports, not a single delivery at the end.

  4. Receive your outputs

    You get the compressed checkpoint, plan file, and evaluation report.

Evidence

Measured result: ViT-B/16

ViT-B/16 · ImageNet-1k · 34% uniform · one seed

Parameters (M)

0 20 40 60 80 100 86.6 30.5 2.84× fewer

Accuracy %

0 20 40 60 80 100 84.5 77.6 ViT-B/16
Dense baseline Compressed 84.51% → 77.58% · 86.6M → 30.5M (2.84×)
34% uniform · one seed. No latency or energy measurements reported. Other architectures and densities: measured on your evaluation.

Request a proposal

We review manually and send a proposal.

Submissions go to info@primeneural.net.