Blackhole

Tenstorrent’s Blackhole is a family of PCIe AI accelerator cards designed for developers and researchers who need auditable, open-source hardware for inference and small-scale training.

Reviewed by 7wData

On this page

Publisher review

Tenstorrent’s Blackhole is a family of PCIe AI accelerator cards designed for developers and researchers who need auditable, open-source hardware for inference and small-scale training. Unlike NVIDIA’s proprietary stack, Blackhole uses a fully open RISC-V instruction set architecture (ISA) and an MIT-licensed compiler stack (TT-Metalium, TT-NN, TT-Forge) that you can read, modify, and submit pull requests to. The cards target workloads where data-movement efficiency and full-stack transparency matter more than raw peak FLOPS — for example, sovereign AI programs, defense and financial services compliance, and hardware research groups. The lineup includes three models: the p100a (active-cooled, no Ethernet, $999), p150a (active-cooled with four passive QSFP-DD 800G ports, $1,399), and p150b (passive-cooled, $1,399). A liquid-cooled desktop workstation, the TT-QuietBox, packs four Blackhole ASICs for $11,999.

Each Blackhole chip integrates 120 Tensix Cores — self-contained compute tiles combining a RISC-V data-movement processor, an 8x8 BF16 matrix engine, and a vector unit — with 1.5 MB of local L1 SRAM per core, yielding 180 MB of on-chip SRAM aggregate. The p100a carries 28 GB of GDDR6 on a 256-bit bus delivering 448 GB/s bandwidth; the p150a and p150b carry 32 GB of GDDR6 at 512 GB/s. Built on a 6nm manufacturing process, the cards draw up to 300W. The p150a’s four QSFP-DD 800G ports enable direct chip-to-chip Ethernet fabric (100 Gbps per link) without InfiniBand or NVLink, scaling from a single card to a 32-chip mesh network. The open-source software stack includes TT-Metalium for kernel development (explicit DMA transfers, tile pipelines) and TT-NN for neural network operators.

Blackhole competes directly with NVIDIA’s RTX 5090 and AMD’s Instinct series, but from a fundamentally different architectural philosophy. NVIDIA’s moat — CUDA, cuDNN, cuBLAS, NVLink, TensorRT — is 20 years deep and closed. Blackhole’s RISC-V ISA and open compiler stack offer full auditability, which is a first-class requirement under the EU AI Act and certain national AI programs. However, the software stack is immature: TT-Metalium kernel development is closer to writing FPGA logic than CUDA, and the ecosystem of pre-optimized models and libraries is thin compared to CUDA’s. In terms of raw memory bandwidth, a single p150a (512 GB/s) trails an H100 SXM5 (3.35 TB/s) by roughly 6x, though the large on-chip SRAM (180 MB) can compensate for small-batch inference where the working set fits locally.

The honest trade-offs are clear. Blackhole offers lower cost per card ($999–$1,399) and full transparency, but at the expense of VRAM capacity (28–32 GB vs. 80 GB on H100), memory bandwidth (448–512 GB/s vs. 3.35 TB/s), and software maturity. The explicit data-movement model makes performance predictable when tuned correctly, but debugging is painful — you must think in terms of tile pipelines and explicit DMA transfers. For teams that need auditable hardware and are willing to invest in compiler-level optimization, Blackhole is a viable alternative. For anyone who just wants to run PyTorch models out of the box with minimal friction, the NVIDIA ecosystem remains the safer bet.

Get the AI & data signal, daily.

48k+ subscribers read this every morning. One email, both newsletters. Unsubscribe anytime.

How it works

  1. RISC-V ISA

    Fully open instruction set architecture built on RISC-V, enabling full-stack auditability and modification without proprietary microcode.

  2. 120 Tensix Cores

    Each Blackhole chip packs 120 self-contained compute tiles, each with a RISC-V processor, 8x8 BF16 matrix engine, and vector unit.

  3. Up to 32 GB GDDR6

    p150a and p150b offer 32 GB of GDDR6 on a 256-bit bus, delivering 512 GB/s memory bandwidth per chip.

  4. Mesh-network chip fabric

    Four QSFP-DD 800G ports per p150a enable direct chip-to-chip Ethernet links at 100 Gbps each, scaling to 32 chips without InfiniBand.

  5. Open-source software stack

    MIT-licensed TT-Metalium, TT-NN, TT-Forge, and TT-LLK stacks on GitHub allow reading, modifying, and submitting pull requests to every kernel.

  6. 6nm manufacturing

    Built on a 6nm process, Blackhole cards operate at up to 300W with active or passive cooling options.

  7. 180 MB on-chip SRAM

    Aggregate 1.5 MB L1 SRAM per Tensix Core across 120 cores provides 180 MB of fast local memory for tiled inference workloads.

Strengths and trade-offs

Strengths

  • Fully open RISC-V ISA and MIT-licensed compiler stack enable full-stack auditability, meeting EU AI Act and sovereign AI compliance requirements.
  • Mesh-network chip fabric with four QSFP-DD 800G ports per p150a allows scaling from one card to 32 chips without proprietary interconnects like NVLink or InfiniBand.
  • Lower cost per card ($999–$1,399) compared to NVIDIA RTX 5090 or AMD Instinct equivalents, with a liquid-cooled 4-chip workstation at $11,999.
  • 180 MB of on-chip SRAM across 120 Tensix Cores provides fast local memory for small-batch inference, reducing reliance on GDDR6 bandwidth.

Trade-offs

  • Software stack is immature: TT-Metalium kernel development requires explicit DMA transfers and tile pipelines, closer to FPGA design than CUDA, with a thin ecosystem of pre-optimized models.
  • VRAM capacity tops out at 32 GB per chip, significantly less than NVIDIA H100's 80 GB HBM3, limiting large-batch inference and large KV cache workloads.
  • Memory bandwidth of 512 GB/s (p150a) trails H100's 3.35 TB/s by roughly 6x, making GDDR6 a bottleneck when on-chip SRAM is insufficient.
  • Higher price per GB of VRAM compared to competitors: $1,399 for 32 GB ($43.7/GB) versus RTX 5090's $1,999 for 32 GB ($62.5/GB) but with less mature software support.

Pricing context

p100a (active-cooled, no Ethernet): $999; p150a (active-cooled, 4x QSFP-DD 800G): $1,399; p150b (passive-cooled): $1,399; QSFP-DD 800G cable: $200; TT-QuietBox (liquid-cooled, 4 ASICs): $11,999.

Getting started with Blackhole

  1. Order a Blackhole card

    Visit Tenstorrent's website and select the Blackhole model that fits your workload: p100a ($999) for single-card inference, or p150a/p150b ($1,399) if you need Ethernet mesh scaling. Add the card to your cart and complete the purchase.

  2. Install the card physically

    Power down your system, open the chassis, and insert the Blackhole PCIe card into an available x16 slot. Secure it with the bracket, connect power cables (up to 300W), and close the system. For p150a, attach QSFP-DD cables if networking.

  3. Install the open-source stack

    Clone the TT-Metalium and TT-NN repositories from GitHub. Follow the README to install dependencies (Python, CMake, GCC) and build the compiler stack. Run the provided smoke tests to verify the card is detected and functional.

  4. Write a kernel for inference

    Using TT-Metalium, define a tile pipeline that loads weights from GDDR6 into Tensix Core L1 SRAM, performs matrix multiplication on the 8x8 BF16 engine, and writes results back. Compile and run the kernel on a single chip to validate correctness.

  5. Scale to multi-chip mesh

    If using p150a cards, connect them via QSFP-DD 800G cables to form a mesh network. Configure TT-NN to distribute tensor operations across chips using explicit DMA transfers. Run a small batch inference to verify inter-chip communication.

Frequently Asked Questions

What is Tenstorrent Blackhole and who is it for?

Tenstorrent Blackhole is a family of PCIe AI accelerator cards using an open RISC-V instruction set and MIT-licensed compiler stack. It targets developers and researchers who need auditable hardware for inference and small-scale training, especially in sovereign AI, defense, and financial services.

How does Blackhole's open-source software stack work?

Blackhole's software stack includes TT-Metalium for kernel development, TT-NN for neural network operators, and TT-Forge, all under an MIT license. Developers can read, modify, and submit pull requests to every kernel, enabling full-stack transparency and auditability.

What are the specs and pricing for Blackhole models?

The p100a has 28 GB GDDR6, 448 GB/s bandwidth, and costs $999. The p150a and p150b have 32 GB GDDR6, 512 GB/s bandwidth, and cost $1,399 each. The p150a adds four QSFP-DD 800G ports for mesh networking.

How does Blackhole compare to NVIDIA H100 in memory and performance?

A single Blackhole p150a offers 32 GB GDDR6 with 512 GB/s bandwidth, while an H100 SXM5 has 80 GB HBM3 at 3.35 TB/s. Blackhole's 180 MB on-chip SRAM helps for small-batch inference, but it trails in large-batch workloads.

Can Blackhole scale to multiple chips without InfiniBand?

Yes, the p150a model has four QSFP-DD 800G ports that enable direct chip-to-chip Ethernet links at 100 Gbps each. This allows scaling from a single card to a 32-chip mesh network without proprietary interconnects like NVLink or InfiniBand.

What are the main trade-offs of using Blackhole over NVIDIA?

Blackhole offers lower cost per card and full auditability but has immature software requiring explicit DMA and tile pipeline programming. It also has less VRAM and memory bandwidth than H100, making it better for teams willing to invest in compiler-level optimization.

Alternatives

How Blackhole compares

Direct head-to-head against 3 competitors. Picked by 7wData.

This tool

Blackhole

Pricing
p100a (active-cooled, no Ethernet): $999; p150a (active-cooled, 4x QSFP-DD 800G): $1,399; p150b (passive-cooled): $1,399; QSFP-DD 800G cable: $200; TT-QuietBox (liquid-cooled, 4 ASICs): $11,999.
Target
Tenstorrent’s Blackhole is a family of PCIe AI accelerator cards designed for developers and researchers who need auditable, open-source hardware for inference and small-scale training.
Strength
Fully open RISC-V ISA and MIT-licensed compiler stack enable full-stack auditability, meeting EU AI Act and sovereign AI compliance requirements.
Watch for
Software stack is immature: TT-Metalium kernel development requires explicit DMA transfers and tile pipelines, closer to FPGA design than CUDA, with a thin ecosystem of pre-optimized models.

VideoLan

Pricing
Free
Target
Media playback and streaming
Deployment
Desktop
Strength
Broad format support
Watch for
Limited advanced features

Konvey

Pricing
Custom/Contact sales
Target
Screen recording and branding
Deployment
Desktop
Strength
Integrated branding tools
Watch for
Complex setup

Handbrake

Pricing
Free
Target
Video conversion
Deployment
Desktop
Strength
High-quality video conversion
Watch for
Steep learning curve

User reviews

No user reviews yet. Be the first to write one.

Sources

Reporting on this tool draws on these publicly available sources.

  1. www.reddit.com
  2. tenstorrent.com
  3. news.ycombinator.com
  4. www.spheron.network