TileLang explained: write AI kernels at tile granularity
TileLang performance and cost: benchmark the whole kernel lifecycle
Separate compilation, warm execution, correctness and service-level value
What you will learn
- Make the reference fair
- Measure cold and warm paths
- Price engineering time
Before you start
- One target operator and device
- A correct framework reference
- A compatible compiler and runtime
Move from a passing README example to evidence for one real model bottleneck
Key takeaways
- Upstream tables are not local measurements.
- Rare shapes can make JIT cost visible.
- Operator speed is not automatically model speed.
Make the reference fair
Compare a TileLang operator with a correct implementation for the same shape, dtype, layout and hardware. Confirm output tolerances before timing; attention workloads need equally precise masking and sequence-length conditions.
The upstream benchmark summary describes selected results, but this review did not reproduce them. Treat those values as project claims with their own scripts and environment, not as numbers for your deployment.
Measure cold and warm paths
Report import, first-call compilation, warm kernel latency, end-to-end model latency and peak memory separately. A fast warm kernel may be unsuitable if the service sees many rare shapes that continually compile.
Autotuning can search tile choices at additional compile and measurement cost. Fix the search space and run count, then save selected configurations so the result can be explained and repeated.
Price engineering time
A hand-written kernel adds validation across shapes, toolchains and new devices. A small operator gain may disappear behind framework overhead or model bottlenecks elsewhere.
Decide on adoption using an accepted model-level objective, not a single microbenchmark. No GPU benchmark or cost study was run for this series.
Decision guide
| Criterion | Option A | Option B |
|---|---|---|
| Best when | You need predictable behavior and easy auditing | You need adaptive optimization and have reliable telemetry |
| Main risk | May leave performance on the table | Can become difficult to explain or debug |
Implementation steps
- 1
Match shape, dtype, layout and backend in the reference.
- 2
Report compile, warm, model and memory metrics separately.
- 3
Count autotuning and maintenance effort.
Copy-ready example
shape,dtype,device,reference_ok,compile_ms,warm_ms,model_ms,peak_mb
,,,,not-run,,,Frequently asked questions
What is the first useful metric?
A correct result for the target shape and device, followed by separate cold and warm timings.
Will autotuning always improve the service?
Not necessarily; its search costs and chosen specialization must fit the service workload.
Sources
- TileLang / README.mdSource checked 2026-10-04
- TileLang / docs/tutorials/auto_tuning.mdSource checked 2026-10-04
- TileLang / examples/gemm/README.mdSource checked 2026-10-04
- TileLang / examples/flash_attention/README.mdSource checked 2026-10-04
- TileLang / docs/get_started/Installation.mdSource checked 2026-10-04