Bonsai Demo explained: local inference with a fork-specific model
Bonsai Demo explained: local inference with a fork-specific model
Separate the demo repository, downloaded model weights and required runtime before judging the claims.
What you will learn
- Name the three artifacts
- Understand the default route
- Keep the test bounded
Before you start
- A target machine with measured memory and disk
- Permission to inspect downloaded model and binary files
Capture the exact artifacts and one correctness failure before claiming a useful deployment.
Key takeaways
- The demo, weights and runtime are distinct artifacts.
- Bonsai 2 currently requires the PrismML fork.
- Upstream quality figures are not results of this review.
Name the three artifacts
Bonsai Demo provides setup and launch scripts, optional UI and agent examples, and documentation for Bonsai 2 27B. Model weights are downloaded from a separate Hugging Face repository. The demo also fetches PrismML’s llama.cpp fork because stock llama.cpp does not currently implement the model’s required transforms.
The README describes a compact ternary model and publishes quality figures. Those are upstream claims tied to its tests. This series inspects fixed source files but has not downloaded the weights or measured model quality.
Understand the default route
At the inspected revision, plain `./setup.sh` selects Bonsai 2 and the PQ2_0 packing, then `start_llama_server.sh` starts a local chat endpoint. Installation and starting the server are separate actions, which makes it easier to inspect downloads first.
The README also offers MLX on Apple Silicon, Open WebUI, tool calls and vision. Do not assume every backend supports every packing or feature; the backend matrix and exact model file determine the path.
Keep the test bounded
A useful first trial asks a few fixed text questions, notes the model file, binary revision and hardware, then checks that the responses are coherent. A server returning HTTP 200 does not prove the required model transforms are active.
The next chapters cover setup, format compatibility, source scripts, performance evaluation, safety and selection. None of the speed or quality numbers in the upstream README are presented as our benchmark.
Decision guide
| Criterion | Option A | Option B |
|---|---|---|
| Best when | You need predictable behavior and easy auditing | You need adaptive optimization and have reliable telemetry |
| Main risk | May leave performance on the table | Can become difficult to explain or debug |
Implementation steps
- 1
Identify demo code, model file and fork binary separately.
- 2
Read the backend support matrix for your device.
- 3
Plan a small correctness check before performance measurement.
Copy-ready example
demo scripts -> fork-specific llama.cpp binary
model repository -> PQ2_0 or PTQ1_0 weights
binary + weights -> local inference -> correctness checkFrequently asked questions
Can stock llama.cpp run Bonsai 2?
The fixed README and support matrix say to use the PrismML fork for all Bonsai 2 GGUF variants.
Are weights included in this GitHub repository?
No. The setup downloads model files from a separate distribution.
Sources
- Bonsai Demo / README.mdSource checked 2026-09-29
- Bonsai Demo / MODEL-FORMATS.mdSource checked 2026-09-29
- Bonsai Demo / BACKEND-SUPPORT.mdSource checked 2026-09-29