Developers

Get your models up and running fast on Tenstorrent hardware. With two open source SDKs, you can get as close to the metal as possible, or let our AI compiler do the work.

Models

Explore models optimized on Tenstorrent hardware.

Don’t see yours listed? Check out TT-Forge, our compiler, to get other models running today.

Models48
AFM-4.5B

Compact foundation model tuned for efficient edge and on-device deployment. Arcee AI, 4.5B.

Text Generation
4.5B
bge-large-en-v1.5

Purpose-built for semantic search, dense retrieval, and RAG. BAAI, 335M.

Feature Extraction
335M
BGE-M3

Multi-lingual, multi-granularity retrieval across dense, sparse, and multi-vector modes. BAAI, 568M.

Feature Extraction
568M
DeepSeek-R1-0528

Refreshed R1 checkpoint - mixture-of-experts reasoning with a fraction active per token. DeepSeek, 671B.

Text Generation
671B
DeepSeek-R1

Mixture-of-experts reasoning powerhouse rivaling closed frontier models on math and code. DeepSeek, 671B.

Text Generation
671B
DeepSeek-R1-Distill-Llama-70B

R1 reasoning traces distilled into a Llama 70B backbone for math and code. DeepSeek, 70B.

Text Generation
70B
diffusiongemma-26B-A4B-it

Instruction-tuned text-diffusion Gemma 4 MoE that writes whole 256-token blocks per step, 4B active of 26B total. Google, 26B.

Text Generation
26B
distil-large-v3

Distilled Whisper large-v3 — near-parity transcription at roughly six times the speed. Distil-Whisper, 756M.

Speech-to-Text
756M
Falcon3-7B-Instruct

Third-generation Falcon instruction model with a 32K context window. TII, 7B.

Text Generation
7B
FLUX.1 [dev]

Flow-matching transformer for photorealistic text-to-image with strong prompt adherence. Black Forest Labs, 12B.

Text-to-Image
12B
FLUX.1 [schnell]

FLUX distilled to 4 steps — full quality, fraction of the compute. Black Forest Labs, 12B.

Text-to-Image
12B
gemma-3-27b-it

Largest Gemma 3 — 140-language reasoning, coding, long context. Google, 27B.

Text Generation
27B
gemma-4-31b-it

Instruction-tuned Gemma 4 for multilingual reasoning, coding, and long context. Google, 31B.

Text Generation
31B
gpt-oss-120b

Open GPT-style model for deep reasoning in self-hosted deployments. GPT-OSS, 120B.

Text Generation
120B
Llama-3.1-70B

128K context, multilingual, tool-use ready, fully open weights. Meta, 70B.

Text Generation
70B
Llama-3.1-70B-Instruct

Complex reasoning, coding, and agentic pipelines — no license restrictions. Meta, 70B.

Text Generation
70B
Llama-3.1-8B

128K context and multilingual base with a wide fine-tuning ecosystem. Meta, 8B.

Text Generation
8B
Llama-3.1-8B-Instruct

Multilingual instruction following, tool use, and function calling. Meta, 8B.

Text Generation
8B
Llama-3.2-1B

Sub-gigabyte base model for on-device and embedded inference. Meta, 1B.

Text Generation
1B
Llama-3.2-1B-Instruct

Instruction-tuned to run anywhere — on-device at 1B. Meta, 1B.

Text Generation
1B
Llama-3.2-3B

On-device reasoning with room for language understanding and light tool use. Meta, 3B.

Text Generation
3B
Llama-3.2-3B-Instruct

Lightweight agent backbone for low-latency, resource-constrained deployments. Meta, 3B.

Text Generation
3B
Llama-3.3-70B-Instruct

Stronger structured tasks, tool use, and reasoning than prior Llama generations. Meta, 70B.

Text Generation
70B
Mistral-7B-Instruct-v0.3

Function-calling and instruction following for production and agentic use. Mistral AI, 7B.

Text Generation
7B
MobileNet V2

Inverted residuals and linear bottlenecks — strong accuracy at near-zero inference cost. Google, 3.4M.

Image Classification
3.4M
Mochi 1

Text-to-video focused on motion quality and temporal coherence. Genmo, 10B.

Text-to-Video
10B
Qwen2.5-72B

Chinese, English, code, and structured output — 128K context base. Alibaba, 72B.

Text Generation
72B
Qwen2.5-72B-Instruct

Instruction-tuned across Chinese, English, coding, and structured output at scale. Alibaba, 72B.

Text Generation
72B
Qwen3-32B

Toggleable chain-of-thought for on-demand deep reasoning. Alibaba, 32B.

Text Generation
32B
Qwen3.6-27B

Mid-scale Qwen3.6 for general reasoning, coding, and multilingual generation. Alibaba, 27B.

Text Generation
27B
Qwen3-8B

Reasoning depth on demand without a throughput penalty. Alibaba, 8B.

Text Generation
8B
Qwen3-Embedding-0.6B

Compact multilingual embeddings for retrieval on constrained hardware. Alibaba, 0.6B.

Embedding
0.6B
Qwen3-Embedding-4B

Multilingual embeddings for retrieval and semantic similarity. Alibaba, 4B.

Embedding
4B
Qwen3-Embedding-8B

Higher-capacity multilingual embeddings for retrieval and reranking. Alibaba, 8B.

Embedding
8B
Qwen3-VL-32B-Instruct

Hybrid thinking + vision — documents, charts, UI, and multi-image tasks. Alibaba, 32B.

Image-Text-to-Text
32B
QwQ-32B

Chain-of-thought with self-reflection for math, science, and logic. Alibaba, 32B.

Text Generation
32B
ResNet-50

Skip connections that made deep networks trainable — the image classification baseline. Microsoft Research, 25M.

Image Classification
25M
SegFormer (b0)

Mix-transformer segmentation without positional encoding - accurate at low compute. NVIDIA, 3.8M.

Image Segmentation
3.8M
SpeechT5 (TTS task)

Unified encoder-decoder for natural text-to-speech synthesis. Microsoft, 307M.

Text-to-Speech
307M
Stable Diffusion 3.5 Large

MMDiT architecture — strong text rendering and prompt control for image generation. Stability AI, 8B.

Text-to-Image
8B
SD-XL 1.0-base

Dual-encoder base for high-resolution synthesis and fine-tuning pipelines. Stability AI, 6.6B.

Text-to-Image
6.6B
SD-XL 1.0-base (img2img)

SD-XL base driven image-to-image - edits or restyles an existing image from a prompt. Stability AI, 6.6B.

Text-to-Image
6.6B
vit-base

Image patches + self-attention, no convolutions — the original vision transformer. Google, 86M.

Image Classification
86M
vovnet-19b-ra

One-Shot Aggregation avoids DenseNet's redundant paths for better accuracy-per-FLOP. 11.2M.

Image Classification
11.2M
Wan2.2

Causal video transformer for text-to-video with strong motion coherence. Alibaba, 14B.

Text-to-Video
14B
whisper-large-v3

99 languages, 680K hours of training audio — built for robust speech recognition. OpenAI, 1.5B.

Speech-to-Text
1.5B
YOLOX-Nano

Anchor-free detector tuned for real-time inference at minimal parameter cost. Megvii, 0.9M.

Object Detection
0.9M
Z-Image-Turbo

Few-step distilled text-to-image generation. Tongyi-MAI.

Text-to-Image
6B

Start a bounty

Build an open future with us. Fix bugs, add features, get paid.

Join our developer community

Get the latest info, ask questions, review our open-source repos.