Developers
Get your models up and running fast on Tenstorrent hardware. With two open source SDKs, you can get as close to the metal as possible, or let our AI compiler do the work.
Models
Explore models optimized on Tenstorrent hardware.
Don’t see yours listed? Check out TT-Forge, our compiler, to get other models running today.
Models47
Sep 4Sep 4Sep 2Aug 27Aug 15Aug 6Aug 5Aug 4Aug 3Jul 28
Fix INT_MIN correctness in int32 div, remainder, fmod, and scalar promotion
hard
DRAM-sharded decode matmul: the in0 multicast is one hop per activation shard
hard
binary_ng/ternary: replicate the broadcast row with local NoC copies instead of ~2000 scalar stores per tile
medium
ttnn.prod_bw returns non-finite gradients for zero inputs
medium
ttnn.softplus takes the linear branch at exactly input*beta == threshold (strict < where reference uses <=)
medium
Generalize multi_scale_deformable_attn to support D values that are multiples of 16
medium
Fix Blackhole destination-reuse synchronization
hard
ttnn/normalization: sharded softmax accepts a subblock_w larger than Dest and silently returns wrong results
medium
Fix SFPU RNG correlation and improve FP32 uniform random quality
hard
Gemma-2 (2B / 9B) text model bring-up using TTNN APIs
hard