개발자
Tenstorrent 하드웨어에서 모델을 빠르게 실행하세요. 두 개의 오픈 소스 SDK로 가능한 한 금속에 가까이 접근하거나 AI 컴파일러에 맡길 수 있습니다.
모델
Explore models optimized on Tenstorrent hardware.
Don’t see yours listed? Check out TT-Forge, our compiler, to get other models running today.
모델46
Sep 16Sep 4Sep 4Sep 2Aug 27Aug 15Aug 12Aug 6Aug 5Aug 4
SFPLOADMACRO produces wrong results under coverage instrumentation (Reciprocal fails 46/153 on Blackhole)
medium
Fix INT_MIN correctness in int32 div, remainder, fmod, and scalar promotion
hard
DRAM-sharded decode matmul: the in0 multicast is one hop per activation shard
hard
binary_ng/ternary: replicate the broadcast row with local NoC copies instead of ~2000 scalar stores per tile
medium
ttnn.prod_bw returns non-finite gradients for zero inputs
medium
ttnn.softplus takes the linear branch at exactly input*beta == threshold (strict < where reference uses <=)
medium
Perf - top-k/sort/moe/sampling index tile is written with 1024 scalar stores per tile on a DM core
hard
Generalize multi_scale_deformable_attn to support D values that are multiples of 16
medium
Fix Blackhole destination-reuse synchronization
hard
ttnn/normalization: sharded softmax accepts a subblock_w larger than Dest and silently returns wrong results
medium