Packages Qwen3.6-35B-A3B (Q5_K_XL GGUF), a matching runtime
(Intel SYCL, NVIDIA CUDA, or portable CPU), and local
llama-server with /v1/chat/completions on loopback.
Numbers on this page name their test and link to a result file
(claims rule).
Clean upstream SYCL binary versus Treebeard package binary with package env.
Same weights: Q5_K_XL sha256 25233af7…c506.
| Test | Shape | Control | Treebeard | Artifact |
|---|---|---|---|---|
| tool-eval-bench public 69 | np=1 · c=262144 · temp 0 · seed 42 · no-think · 2.1.0@8b3259b | 91/100 (125/138) | 91/100 (126/138) | agent-bench-ab REPORT |
| held-out ho-pack-v1.1 | np=1 · temp 0 · seed 42 · 23 gated scenarios | 42/46 (91.3%) | 42/46 (91.3%) | same dir · heldout/*/result.json |
| single-agent sequential smoke | np=1 · c=32768 · 5 prompts × 2 · tg_p50 | 77.1 tok/s | 88.9 tok/s (+15.4%) | single-agent-ab REPORT |
| 12-agent concurrent ABA | np=12 · n_predict=96 · temp 0 · p50 tok/s per agent | 6.88 | 26.33 (+282.8%) | base-vs-package-aba REPORT |
| 12-agent aggregate | same ABA · total tok/s | 72.9 | 208.7 | same REPORT |
Multi-slot p50 is concurrent capacity. Single-user chat speed is the sequential row. Public suite status differed on TC-50 only (control partial, Treebeard pass). Charts: dual-axis, quality A/B, multi-slot. Writeup: RELEASE-20260728.md.
| Treebeard-only run (no control arm) | Result | Artifact |
|---|---|---|
| Broadway multi-page HTML structure (6 sites) | 6/6 | broadway-eval REPORT |
| Long-context needles (3 depths) | 3/3 | longctx-eval REPORT |
| Long-context dossier QA | 14/14 | same longctx REPORT |
| Long multi-slot stress (c=262144, np=12) | 120/120 request success | long-multislot-stress REPORT |
Checksum-pinned package run on Intel Arc Pro B70, with an independent replica on NVIDIA GB10. Shape: np=1, c=262144, temp 0, thinking off, seed 42. The 2026-07-28 stock-Q5 control A/B above scored 91 on both arms under a different pin and quant context.
Category bars are read from the B70 freeze result file. Methodology: package/docs/BENCHMARKS.md.
| Test | Hardware | Result | Artifact |
|---|---|---|---|
| tool-eval-bench 69 control A/B | B70 · stock Q5 · np=1 | 91 = 91 | private-verification…/agent-bench-ab-… |
| ho-pack-v1.1 control A/B | B70 · stock Q5 · np=1 | 42/46 = 42/46 | same · heldout/ |
| single-agent sequential tg_p50 | B70 · stock Q5 · np=1 · c=32768 | 77.1 → 88.9 | private-verification…/single-agent-ab-… |
| 12-agent ABA p50/agent | B70 · stock Q5 · np=12 | 6.88 → 26.33 | private-verification…/base-vs-package-aba-… |
| tool-eval-bench 69 freeze | B70 + GB10 | 94/100 both | results/agent/single-slot-94/ |
| llama-bench pp4096 | GB10 CUDA | 2,026.9 tok/s | results/nvidia/native-bench/ |
| llama-bench tg128 | GB10 CUDA | 52.5 tok/s | same |
| Q8_0 12-col latency | Blackwell CUDA | 33.0% lower (direct path) | results/nvidia/attribution-q8/ |
| Q8_0 MoE-down latency | Blackwell CUDA | 3.4% lower | same |
| Installed-package chat smoke | Ryzen 9 5950X CPU | 9.30 tok/s chat; exact tool call | results/cpu-linux-x86_64/smoke/ |
| 12-slot aggregate serving | B70 SYCL ship profile | 194.023 tok/s aggregate | results/sycl/ |
agent-bench-ab-…heldout/single-agent-ab-…base-vs-package-aba-…results/agent/single-slot-94/.TREEBEARD_REASONING default off (freeze shape: thinking disabled).TREEBEARD_SPECULATION opt-in; no speed claim on this page.
Download is about 26.7 GB (resumable). Installer checks file SHA-256.
Plan on roughly 32 GB system or unified memory.
--multimodal adds the 0.9 GB vision projector.
curl -fsSL https://raw.githubusercontent.com/newjordan/treebeard/main/install.sh | bash
Fetches the model and the runtime for this host (SYCL, CUDA, or CPU).
treebeard doctor
Prints platform selection and the launch command.
treebeard serve
Serves http://127.0.0.1:8093/v1.