reap-lab
Shrink big AI models, then prove they still work.
Large open models contain many “specialist” experts a given workload never uses. reap-lab finds and removes the idle ones so the model fits on an ordinary GPU, then re-tests it — and refuses to install it if quality drops.
- Tests passing
- 563
- Minimum quality kept
- 95%
→Measures which experts your own workload activates
→Prunes and compresses (e.g. ~61 GB down to ~15 GB, estimated)
→Quality gate: at least 95% of the original or it is not installed
Two ways to read it
In plain English
Large open models contain many “specialist” experts a given workload never uses. reap-lab finds and removes the idle ones so the model fits on an ordinary GPU, then re-tests it — and refuses to install it if quality drops.
For engineers
A Python 3.11 pipeline that profiles expert activation on your own workload, prunes and compresses the MoE, and gates installation on scoring ≥95% of the original.
What it does.
- Measures which experts your own workload activates
- Prunes and compresses (e.g. ~61 GB down to ~15 GB, estimated)
- Quality gate: at least 95% of the original or it is not installed
Built with
- Python 3.11
- uv
- PyTorch tooling
- GitHub Actions
Status: Shipped · started July 2026 · open on GitHub