TileLang explained: write AI kernels at tile granularity
TileLang quickstart: a GEMM kernel with a correctness gate
Set up one supported device and verify output before tuning tile sizes
What you will learn
- Select an installation route
- Run the smallest example
- Know what a green assertion means
Before you start
- One target operator and device
- A correct framework reference
- A compatible compiler and runtime
Move from a passing README example to evidence for one real model bottleneck
Key takeaways
- ROCm requires host hipcc and ROCm PyTorch.
- `@tilelang.jit` specializes on first use.
- One passing shape is not exhaustive validation.
Select an installation route
For a supported environment the pinned guide starts with `pip install tilelang`. NVIDIA hosts need a usable CUDA toolchain; AMD Linux wheels include ROCm support but still need host `hipcc` and a ROCm PyTorch installed first.
Do not interchange the CUDA `tilelang[nvcc]` extra with a ROCm setup. Record the Python, PyTorch, toolkit, GPU and TileLang versions before comparing results or opening an issue.
Run the smallest example
The README defines `matmul_relu` under `@tilelang.jit`, creates CUDA-device PyTorch tensors, calls the kernel and uses `torch.testing.assert_close` against `torch.relu(a @ b)`.
Start with the documented shapes and tolerances, then vary one dimension at a time. First use includes specialization and compilation, while later calls may reuse compiled work; those timings should never be mixed.
Know what a green assertion means
A match for one shape and one device is a local correctness result, not a proof for all inputs, dtypes or backends. Exercise edge sizes, numeric tolerance and repeated calls before integrating the kernel into a model.
The code shown below is a reading aid taken from the documented flow, not an executed result from this editorial run.
Decision guide
| Criterion | Option A | Option B |
|---|---|---|
| Best when | You need predictable behavior and easy auditing | You need adaptive optimization and have reliable telemetry |
| Main risk | May leave performance on the table | Can become difficult to explain or debug |
Implementation steps
- 1
Install a compatible device runtime and framework build.
- 2
Run the pinned GEMM + ReLU example with its reference assertion.
- 3
Expand shapes and dtypes before measuring speed.
Copy-ready example
import torch, tilelang
# Run the pinned README matmul_relu definition first.
c = matmul_relu(a, b)
torch.testing.assert_close(c, torch.relu(a @ b), rtol=1e-2, atol=1e-2)Frequently asked questions
Does `pip install tilelang` alone make AMD ready?
No. The guide requires a host ROCm installation and a matching ROCm build of PyTorch first.
Should I time the first and second call together?
No. The first call may compile a specialization; report cold and warm timings separately.
Sources
- TileLang / README.mdSource checked 2026-10-04
- TileLang / docs/get_started/Installation.mdSource checked 2026-10-04
- TileLang / pyproject.tomlSource checked 2026-10-04
- TileLang / tilelang/jit/__init__.pySource checked 2026-10-04