Developers
Get your models up and running fast on Tenstorrent hardware. With two open source SDKs, you can get as close to the metal as possible, or let our AI compiler do the work.
Models
Explore models optimized on Tenstorrent hardware.
Don’t see yours listed? Check out TT-Forge, our compiler, to get other models running today.
Models46
Sep 16Sep 4Sep 4Sep 2Aug 27Aug 15Aug 12Aug 6Aug 5Aug 4
SFPLOADMACRO produces wrong results under coverage instrumentation (Reciprocal fails 46/153 on Blackhole)
medium
Fix INT_MIN correctness in int32 div, remainder, fmod, and scalar promotion
hard
DRAM-sharded decode matmul: the in0 multicast is one hop per activation shard
hard
binary_ng/ternary: replicate the broadcast row with local NoC copies instead of ~2000 scalar stores per tile
medium
ttnn.prod_bw returns non-finite gradients for zero inputs
medium
ttnn.softplus takes the linear branch at exactly input*beta == threshold (strict < where reference uses <=)
medium
Perf - top-k/sort/moe/sampling index tile is written with 1024 scalar stores per tile on a DM core
hard
Generalize multi_scale_deformable_attn to support D values that are multiples of 16
medium
Fix Blackhole destination-reuse synchronization
hard
ttnn/normalization: sharded softmax accepts a subblock_w larger than Dest and silently returns wrong results
medium