Bonsai Demo explained: local inference with a fork-specific model
Bonsai architecture: ternary packing, transforms and serving
Understand why file format support alone cannot establish correct inference.
What you will learn
- Follow model data into compute
- Read the backend matrix carefully
- Place optional layers around the core
Before you start
- A target machine with measured memory and disk
- Permission to inspect downloaded model and binary files
Capture the exact artifacts and one correctness failure before claiming a useful deployment.
Key takeaways
- Packing support and model transform support are different.
- Backend tables describe a release, not every device.
- Optional UI and tools increase the system boundary.
Follow model data into compute
MODEL-FORMATS.md distinguishes Bonsai 2 PTQ1_0, PQ2_0 and a development Q2_0 band. The files store weights in a rotated basis and need the activation transform carried by PrismML’s fork.
A backend able to decode a quantized type may still lack the model graph’s Hadamard and sign-flip operations. That is why the support matrix separates format kernels from transform support.
Read the backend matrix carefully
BACKEND-SUPPORT.md records source-level support for CPU, Metal, CUDA, ROCm, Vulkan and SYCL at a particular fork release. Partial cells and fallbacks matter. A directory named `vulkan` does not prove every operation stayed on a GPU.
The demo’s selection registry chooses a path, but the docs say it is not a full capability probe. Check launch arguments, detected devices and a correctness test on your actual hardware.
Place optional layers around the core
The core local flow is setup, compatible runtime, weights and inference. Open WebUI, tools, agent demo and vision add separate dependencies and permissions. Add one layer at a time so failures remain attributable.
This is a source-based architectural map. It does not reconstruct the full training pipeline or audit the separately hosted weight files.
Decision guide
| Criterion | Option A | Option B |
|---|---|---|
| Best when | You need predictable behavior and easy auditing | You need adaptive optimization and have reliable telemetry |
| Main risk | May leave performance on the table | Can become difficult to explain or debug |
Implementation steps
- 1
Trace model filename to format and fork release.
- 2
Check both format and transform support on the backend.
- 3
Add UI or tools only after core text inference works.
Copy-ready example
weights -> packing decoder -> Hadamard/sign-flip transforms
-> model graph -> token generation -> local API
backend check: format AND transform AND deviceFrequently asked questions
Why can Q2_0 produce nonsense on stock llama.cpp?
The docs say the file can load while Bonsai 2 transforms are missing.
Does downloading PQ2_0 prove native Vulkan support?
No. The matrix notes that selection and native backend support are separate questions.
Sources
- Bonsai Demo / MODEL-FORMATS.mdSource checked 2026-09-29
- Bonsai Demo / BACKEND-SUPPORT.mdSource checked 2026-09-29
- Bonsai Demo / README.mdSource checked 2026-09-29