开发者
在Tenstorrent硬件上快速运行您的模型。通过两个开源SDK,您可以尽可能接近底层硬件,或让我们的AI编译器来完成工作。
模型
Explore models optimized on Tenstorrent hardware.
Don’t see yours listed? Check out TT-Forge, our compiler, to get other models running today.
模型46
Sep 16Sep 4Sep 4Sep 2Aug 27Aug 15Aug 12Aug 6Aug 5Aug 4
SFPLOADMACRO produces wrong results under coverage instrumentation (Reciprocal fails 46/153 on Blackhole)
medium
Fix INT_MIN correctness in int32 div, remainder, fmod, and scalar promotion
hard
DRAM-sharded decode matmul: the in0 multicast is one hop per activation shard
hard
binary_ng/ternary: replicate the broadcast row with local NoC copies instead of ~2000 scalar stores per tile
medium
ttnn.prod_bw returns non-finite gradients for zero inputs
medium
ttnn.softplus takes the linear branch at exactly input*beta == threshold (strict < where reference uses <=)
medium
Perf - top-k/sort/moe/sampling index tile is written with 1024 scalar stores per tile on a DM core
hard
Generalize multi_scale_deformable_attn to support D values that are multiples of 16
medium
Fix Blackhole destination-reuse synchronization
hard
ttnn/normalization: sharded softmax accepts a subblock_w larger than Dest and silently returns wrong results
medium