Sohu
By Etched
Sohu is a transformer-specific AI chip developed by Etched AI, designed exclusively for transformer model inference.
Publisher review
Sohu is a transformer-specific AI chip developed by Etched AI, designed exclusively for transformer model inference. Launched in 2022, Sohu targets enterprises and cloud providers seeking high-throughput, cost-efficient AI inference solutions. Built on TSMC’s 4-nanometer process, the chip leverages a custom architecture with specialized cores and memory configurations to optimize transformer workloads.
Its focus on transformer attention as fixed-function silicon distinguishes it from general-purpose GPUs, making it a niche but potentially transformative solution for large-scale AI deployments. Etched AI has raised $500 million at a $5 billion valuation, signaling strong investor confidence in its specialized approach. Sohu is not yet available for purchase or rent, but the company has booked over $1 billion in contracts, with the first racks slated for delivery in summer 2026.
The chip’s architecture is built around the assumption that transformer models will dominate AI workloads for the foreseeable future, making it a high-risk, high-reward bet on the future of AI infrastructure. Sohu’s performance claims, such as delivering 500,000 tokens per second on Llama 70B, position it as a potential disruptor in the AI hardware market. However, its lack of programmability and reliance on transformer-specific workloads introduce significant limitations.
The chip’s design eliminates general-purpose overhead, enabling it to achieve high utilization rates and potentially lower power consumption compared to GPUs. This makes it particularly appealing for enterprises grappling with rising data-center energy costs. Despite its promise, Sohu’s narrow focus means it cannot adapt to changes in transformer architectures or support other AI models.
This inflexibility, combined with the absence of independent benchmarks and public pricing, makes it a speculative investment for many organizations. Etched AI’s strategy hinges on the assumption that transformer models will remain dominant, but this bet could backfire if new architectures emerge. For now, Sohu represents a bold attempt to redefine AI hardware by prioritizing specialization over versatility.
How it works
-
Transformer-specific design
Sohu is optimized exclusively for transformer models, eliminating general-purpose overhead for higher efficiency.
-
High throughput
Delivers over 500,000 tokens per second on Llama 70B, outperforming GPUs in transformer inference.
-
Custom architecture
Uses specialized cores and memory configurations to hardwire transformer attention into silicon.
-
HBM3 memory
Includes 144 GB of HBM3 memory per chip, enabling large batch sizes without performance degradation.
-
Fixed-function silicon
Implements transformer attention as fixed-function logic, bypassing programmable matrix multiply instructions.
-
TSMC 4nm process
Built on TSMC’s advanced 4-nanometer manufacturing process for improved performance and efficiency.
-
Energy efficiency
Claims up to 10x better performance and efficiency than Nvidia GPUs for transformer inference.
Strengths and trade-offs
Strengths
- Sohu achieves over 500,000 tokens per second on Llama 70B, significantly outperforming GPUs in transformer inference.
- The chip’s fixed-function silicon design eliminates general-purpose overhead, reducing power consumption.
- Sohu’s 144 GB of HBM3 memory per chip supports extremely large batch sizes without performance degradation.
- Built on TSMC’s 4-nanometer process, Sohu leverages advanced manufacturing for improved efficiency and performance.
Trade-offs
- Sohu is limited to transformer models and cannot adapt to changes in AI architectures.
- The chip lacks programmability, making it inflexible for future model updates or alternative workloads.
- No independent third-party benchmarks are available to validate Sohu’s performance claims.
- The absence of public pricing and availability makes it a speculative investment for many organizations.
Pricing context
Not publicly available
Getting started with Sohu
-
Request access
Contact Etched AI sales to inquire about Sohu availability and pricing. Provide your organization details and intended use case.
-
Prepare infrastructure
Set up compatible server racks with required power and cooling specifications for Sohu deployment, based on vendor guidelines.
-
Load transformer model
Upload your transformer model weights to the Sohu system, ensuring compatibility with the chip's fixed-function architecture.
-
Configure batch size
Set optimal batch size parameters leveraging Sohu's 144GB HBM3 memory to maximize throughput without performance degradation.
-
Run inference
Execute transformer model inference tasks, monitoring throughput and power consumption metrics against baseline GPU benchmarks.
Frequently Asked Questions
What is the Sohu AI chip?
Sohu is a specialized AI chip designed exclusively for transformer model inference, developed by Etched AI. Launched in 2022, it uses custom architecture with fixed-function silicon to optimize transformer workloads. Built on TSMC’s 4nm process, it targets high-throughput, energy-efficient AI deployments for enterprises and cloud providers.
How does Sohu compare to Nvidia GPUs for AI workloads?
Sohu claims 500,000 tokens per second on Llama 70B, outperforming GPUs in transformer inference. Its fixed-function design eliminates general-purpose overhead, potentially offering 10x better efficiency. However, unlike GPUs, Sohu cannot run non-transformer models or adapt to architectural changes in future AI systems.
What are Sohu's key technical specifications?
Sohu features 144GB HBM3 memory per chip, TSMC’s 4nm manufacturing, and specialized cores hardwired for transformer attention. It bypasses programmable matrix operations with fixed-function silicon. These design choices aim to maximize throughput (500k tokens/sec) while reducing power consumption compared to general-purpose AI accelerators.
When will Sohu be available for purchase?
Sohu isn't currently available for purchase or rent. Etched AI has booked over $1 billion in contracts, with first deliveries scheduled for summer 2026. The company has raised $500 million at a $5 billion valuation, indicating strong market interest despite the product's unproven real-world performance.
What are the main limitations of Sohu's design?
Sohu cannot run non-transformer models or adapt to architectural changes. Its lack of programmability makes it inflexible for future AI developments. No independent benchmarks verify its performance claims, and absent public pricing adds uncertainty. These factors make Sohu a high-risk bet on transformer model dominance.
Why would companies choose Sohu over GPUs?
Enterprises facing high data-center costs may prefer Sohu for its claimed 10x efficiency gains in transformer workloads. Its specialized design enables larger batch sizes without performance drops. However, this comes at the cost of versatility—Sohu only makes sense for organizations fully committed to transformer architectures.
Alternatives
- NVIDIA H100 ↗
- Groq LPU ↗
How Sohu compares
Direct head-to-head against 2 competitors. Picked by 7wData.
Sohu
- Pricing
- Not publicly available
- Target
- Sohu is a transformer-specific AI chip developed by Etched AI, designed exclusively for transformer model inference.
- Strength
- Sohu achieves over 500,000 tokens per second on Llama 70B, significantly outperforming GPUs in transformer inference.
- Watch for
- Sohu is limited to transformer models and cannot adapt to changes in AI architectures.
NVIDIA H100
- Pricing
- $30,000-$40,000 per GPU
- Target
- General-purpose AI training/inference
- Deployment
- Cloud/on-prem
- Strength
- Programmable for any AI workload
- Watch for
- High power draw, supply constraints
Groq LPU
- Pricing
- Custom/Contact sales
- Target
- Low-latency transformer inference
- Deployment
- Cloud/on-prem
- Strength
- Deterministic latency for LLMs
- Watch for
- Limited to compiler-supported models
User reviews
No user reviews yet. Be the first to write one.
Sources
Reporting on this tool draws on these publicly available sources.