Bonsai Demo explained: local inference with a fork-specific model
Bonsai Demo quickstart: inspect downloads before local inference
Use the documented setup route while keeping the model server on loopback.
What you will learn
- Check machine and download size
- Separate install from launch
- Check an answer and the logs
Before you start
- A target machine with measured memory and disk
- Permission to inspect downloaded model and binary files
Capture the exact artifacts and one correctness failure before claiming a useful deployment.
Key takeaways
- Setup downloads large external artifacts.
- The server does not start merely because setup succeeded.
- Coherent output is the first compatibility check.
Check machine and download size
The README’s default setup fetches a Bonsai 2 PQ2_0 model and vision projector, with optional Open WebUI and code interpreter adding more data. Check available disk, RAM and the exact files before starting.
Read setup.sh in the pinned revision. It selects platform binaries and model files, so a convenient one-command installer is still an executable network operation. The series has not run it.
Separate install from launch
The documented macOS/Linux path clones the repository and runs `./setup.sh`, then starts `./scripts/start_llama_server.sh`. For a smaller trial, the README names variables that skip Open WebUI and code interpreter setup.
Use the default loopback host for the first run. Record the selected model file, backend, effective context setting and listening address. Opening the server to a network changes the security problem.
Check an answer and the logs
Ask a deterministic factual question and a deliberately unsupported one. Inspect logs for the loaded packing and device, and save the full response rather than only a screenshot.
A plausible response is a smoke test, not a quality benchmark. If output is gibberish, verify the binary and format pair before changing prompt parameters.
Decision guide
| Criterion | Option A | Option B |
|---|---|---|
| Best when | You need predictable behavior and easy auditing | You need adaptive optimization and have reliable telemetry |
| Main risk | May leave performance on the table | Can become difficult to explain or debug |
Implementation steps
- 1
Read setup.sh and verify disk and memory capacity.
- 2
Install, then start a loopback server separately.
- 3
Record model, binary, logs and two test answers.
Copy-ready example
git clone https://github.com/PrismML-Eng/Bonsai-demo.git
cd Bonsai-demo
BONSAI_OPENWEBUI=0 BONSAI_CODE_INTERPRETER=0 ./setup.sh
./scripts/start_llama_server.shFrequently asked questions
Does setup automatically start the chat server?
The README separates setup from the later start_llama_server.sh command.
Can I skip the optional UI?
The README documents environment variables for skipping Open WebUI and code interpreter setup.
Sources
- Bonsai Demo / README.mdSource checked 2026-09-29
- Bonsai Demo / setup.shSource checked 2026-09-29
- Bonsai Demo / FAQ.mdSource checked 2026-09-29