Bonsai Demo explained: local inference with a fork-specific model
Bonsai source code: trace setup.sh into the launch scripts
Read the distribution path without pretending the model’s training code lives in this demo.
What you will learn
- Inspect the installer decisions
- Follow a local invocation
- Mark where source review stops
Before you start
- A target machine with measured memory and disk
- Permission to inspect downloaded model and binary files
Capture the exact artifacts and one correctness failure before claiming a useful deployment.
Key takeaways
- Setup scripts select external artifacts.
- CLI and server paths need argument comparison.
- The demo source is not the whole inference implementation.
Inspect the installer decisions
At the fixed revision, setup.sh is the entry point for platform detection, binaries, model downloads and optional UI components. Trace the environment variables it reads and the paths it writes before executing it on a workstation.
Record which steps contact external hosts and which files are cached. If a rerun changes artifacts, compare hashes and selected revisions rather than assuming the same command means the same runtime.
Follow a local invocation
scripts/run_llama.sh wraps a terminal prompt path; start_llama_server.sh provides the server path described in README. Compare their selected binary, model and arguments before treating CLI and HTTP outputs as equivalent.
The agent demo script introduces more tooling and permissions. Read its launch boundary and any tool configuration before allowing it to run commands or access network resources.
Mark where source review stops
The actual inference kernels live in a separate fork, and weights live in separate model repositories. This demo’s shell files expose selection and orchestration, not every arithmetic operation in the model.
We did not run the scripts. A runtime trace should save command, environment minus secrets, model hash, process logs and observed response.
Decision guide
| Criterion | Option A | Option B |
|---|---|---|
| Best when | You need predictable behavior and easy auditing | You need adaptive optimization and have reliable telemetry |
| Main risk | May leave performance on the table | Can become difficult to explain or debug |
Implementation steps
- 1
Open setup.sh and list download destinations.
- 2
Compare run_llama.sh with the server launcher.
- 3
Mark the separate fork and model repositories.
Copy-ready example
setup.sh -> choose platform -> fetch fork binary + model
run_llama.sh -> local prompt
start_llama_server.sh -> loopback API
record hashes + logs for an actual testFrequently asked questions
Is the complete model training code in setup.sh?
No. The script orchestrates a demo; model files and runtime are separate.
Did this review instrument the launcher?
No. It inspected the fixed scripts.
Sources
- Bonsai Demo / setup.shSource checked 2026-09-29
- Bonsai Demo / scripts/run_llama.shSource checked 2026-09-29
- Bonsai Demo / scripts/agent/run_agent_demo.shSource checked 2026-09-29
- Bonsai Demo / README.mdSource checked 2026-09-29