Sohu

Sohu is a transformer-specific AI chip developed by Etched AI, designed exclusively for transformer model inference.

Reviewed by 7wData

On this page

Publisher review

Sohu is a transformer-specific AI chip developed by Etched AI, designed exclusively for transformer model inference. Launched in 2022, Sohu targets enterprises and cloud providers seeking high-throughput, cost-efficient AI inference solutions. Built on TSMC’s 4-nanometer process, the chip leverages a custom architecture with specialized cores and memory configurations to optimize transformer workloads.

Its focus on transformer attention as fixed-function silicon distinguishes it from general-purpose GPUs, making it a niche but potentially transformative solution for large-scale AI deployments. Etched AI has raised $500 million at a $5 billion valuation, signaling strong investor confidence in its specialized approach. Sohu is not yet available for purchase or rent, but the company has booked over $1 billion in contracts, with the first racks slated for delivery in summer 2026.

The chip’s architecture is built around the assumption that transformer models will dominate AI workloads for the foreseeable future, making it a high-risk, high-reward bet on the future of AI infrastructure. Sohu’s performance claims, such as delivering 500,000 tokens per second on Llama 70B, position it as a potential disruptor in the AI hardware market. However, its lack of programmability and reliance on transformer-specific workloads introduce significant limitations.

The chip’s design eliminates general-purpose overhead, enabling it to achieve high utilization rates and potentially lower power consumption compared to GPUs. This makes it particularly appealing for enterprises grappling with rising data-center energy costs. Despite its promise, Sohu’s narrow focus means it cannot adapt to changes in transformer architectures or support other AI models.

This inflexibility, combined with the absence of independent benchmarks and public pricing, makes it a speculative investment for many organizations. Etched AI’s strategy hinges on the assumption that transformer models will remain dominant, but this bet could backfire if new architectures emerge. For now, Sohu represents a bold attempt to redefine AI hardware by prioritizing specialization over versatility.

Get the AI & data signal, daily.

335k+ subscribers read this every morning. One email, both newsletters. Unsubscribe anytime.

How it works

  1. Transformer-specific design

    Sohu is optimized exclusively for transformer models, eliminating general-purpose overhead for higher efficiency.

  2. High throughput

    Delivers over 500,000 tokens per second on Llama 70B, outperforming GPUs in transformer inference.

  3. Custom architecture

    Uses specialized cores and memory configurations to hardwire transformer attention into silicon.

  4. HBM3 memory

    Includes 144 GB of HBM3 memory per chip, enabling large batch sizes without performance degradation.

  5. Fixed-function silicon

    Implements transformer attention as fixed-function logic, bypassing programmable matrix multiply instructions.

  6. TSMC 4nm process

    Built on TSMC’s advanced 4-nanometer manufacturing process for improved performance and efficiency.

  7. Energy efficiency

    Claims up to 10x better performance and efficiency than Nvidia GPUs for transformer inference.

Strengths and trade-offs

Strengths

  • Sohu achieves over 500,000 tokens per second on Llama 70B, significantly outperforming GPUs in transformer inference.
  • The chip’s fixed-function silicon design eliminates general-purpose overhead, reducing power consumption.
  • Sohu’s 144 GB of HBM3 memory per chip supports extremely large batch sizes without performance degradation.
  • Built on TSMC’s 4-nanometer process, Sohu leverages advanced manufacturing for improved efficiency and performance.

Trade-offs

  • Sohu is limited to transformer models and cannot adapt to changes in AI architectures.
  • The chip lacks programmability, making it inflexible for future model updates or alternative workloads.
  • No independent third-party benchmarks are available to validate Sohu’s performance claims.
  • The absence of public pricing and availability makes it a speculative investment for many organizations.

Pricing context

Not publicly available

Getting started with Sohu

  1. Request access

    Contact Etched AI sales to inquire about Sohu availability and pricing. Provide your organization details and intended use case.

  2. Prepare infrastructure

    Set up compatible server racks with required power and cooling specifications for Sohu deployment, based on vendor guidelines.

  3. Load transformer model

    Upload your transformer model weights to the Sohu system, ensuring compatibility with the chip's fixed-function architecture.

  4. Configure batch size

    Set optimal batch size parameters leveraging Sohu's 144GB HBM3 memory to maximize throughput without performance degradation.

  5. Run inference

    Execute transformer model inference tasks, monitoring throughput and power consumption metrics against baseline GPU benchmarks.

Frequently Asked Questions

What is the Sohu AI chip?

Sohu is a specialized AI chip designed exclusively for transformer model inference, developed by Etched AI. Launched in 2022, it uses custom architecture with fixed-function silicon to optimize transformer workloads. Built on TSMC’s 4nm process, it targets high-throughput, energy-efficient AI deployments for enterprises and cloud providers.

How does Sohu compare to Nvidia GPUs for AI workloads?

Sohu claims 500,000 tokens per second on Llama 70B, outperforming GPUs in transformer inference. Its fixed-function design eliminates general-purpose overhead, potentially offering 10x better efficiency. However, unlike GPUs, Sohu cannot run non-transformer models or adapt to architectural changes in future AI systems.

What are Sohu's key technical specifications?

Sohu features 144GB HBM3 memory per chip, TSMC’s 4nm manufacturing, and specialized cores hardwired for transformer attention. It bypasses programmable matrix operations with fixed-function silicon. These design choices aim to maximize throughput (500k tokens/sec) while reducing power consumption compared to general-purpose AI accelerators.

When will Sohu be available for purchase?

Sohu isn't currently available for purchase or rent. Etched AI has booked over $1 billion in contracts, with first deliveries scheduled for summer 2026. The company has raised $500 million at a $5 billion valuation, indicating strong market interest despite the product's unproven real-world performance.

What are the main limitations of Sohu's design?

Sohu cannot run non-transformer models or adapt to architectural changes. Its lack of programmability makes it inflexible for future AI developments. No independent benchmarks verify its performance claims, and absent public pricing adds uncertainty. These factors make Sohu a high-risk bet on transformer model dominance.

Why would companies choose Sohu over GPUs?

Enterprises facing high data-center costs may prefer Sohu for its claimed 10x efficiency gains in transformer workloads. Its specialized design enables larger batch sizes without performance drops. However, this comes at the cost of versatility—Sohu only makes sense for organizations fully committed to transformer architectures.

Alternatives

How Sohu compares

Direct head-to-head against 2 competitors. Picked by 7wData.

This tool

Sohu

Pricing
Not publicly available
Target
Sohu is a transformer-specific AI chip developed by Etched AI, designed exclusively for transformer model inference.
Strength
Sohu achieves over 500,000 tokens per second on Llama 70B, significantly outperforming GPUs in transformer inference.
Watch for
Sohu is limited to transformer models and cannot adapt to changes in AI architectures.

NVIDIA H100

Pricing
$30,000-$40,000 per GPU
Target
General-purpose AI training/inference
Deployment
Cloud/on-prem
Strength
Programmable for any AI workload
Watch for
High power draw, supply constraints

Groq LPU

Pricing
Custom/Contact sales
Target
Low-latency transformer inference
Deployment
Cloud/on-prem
Strength
Deterministic latency for LLMs
Watch for
Limited to compiler-supported models

User reviews

No user reviews yet. Be the first to write one.

Sources

Reporting on this tool draws on these publicly available sources.

  1. www.reddit.com
  2. www.reddit.com
  3. www.linkedin.com
  4. www.spheron.network