Featured starting point
TileLang explained: write AI kernels at tile granularity
See where the Python DSL sits between model code and device-specific execution
See where the Python DSL sits between model code and device-specific execution
Featured starting point
See where the Python DSL sits between model code and device-specific execution
Suggested learning path
01 → 09
See where the Python DSL sits between model code and device-specific execution
Set up one supported device and verify output before tuning tile sizes
Move a verified operator into an inference or training service without surprising rebuilds
Trace the DSL, JIT specialization and backend choices without treating one GPU as universal
Use the public example and JIT entry point as an auditable path through the repository
Separate compilation, warm execution, correctness and service-level value
Treat compiler inputs, device runtimes and shape-dependent failures as operational concerns
Pick the least costly route that meets correctness and throughput requirements
Move from a passing README example to evidence for one real model bottleneck