Products

Jul 25, 2026
Difflet: Serving Diffusion Models on AWS Trainium
Difflet runs six production image and video diffusion models end to end on AWS Trainium, with up to 4.10x the throughput per dollar of an H100.

FEATURED POSTS
Products

Jul 25, 2026
Difflet: Serving Diffusion Models on AWS Trainium
Difflet runs six production image and video diffusion models end to end on AWS Trainium, with up to 4.10x the throughput per dollar of an H100.
Inference

Mar 17, 2026
Mini-SGLang-Neuron: Bringing Lightweight LLM Inference to AWS Trainium and Inferentia
A lightweight inference framework integrating SGLang with AWS Neuron to enable efficient LLM serving on Trainium and Inferentia across multi-hardware environments.
Products

Feb 06, 2026
Launch Templates: Infrastructure Portability for Production AI
AI infrastructure should not dictate how you build models.
News

Jan 16, 2026
Yotta Labs Welcomes Jack Dongarra: A Signal for the Next Era of AI Infrastructure
Dr. Jack Dongarra, 2021 ACM A.M. Turing Award recipient and architect of modern performance benchmarking, has joined Yotta Labs as a Technical & Strategic Advisor. As AI infrastructure reaches a new inflection point, Yotta Labs is applying decades of hard-won HPC lessons to build an intelligent orchestration layer for scalable, interoperable GPU systems.
News

Jan 05, 2026
Academic Research Credit Support Program Launch
Artificial intelligence research is advancing at an unprecedented pace — yet access to scalable, reliable compute remains one of the biggest constraints facing researchers today. Across universities, research labs, and independent research communities, ambitious ideas are often slowed by limited GPU availability, high infrastructure costs, and rigid cloud environments not designed for experimentation. Researchers are forced to make tradeoffs: smaller models, fewer experiments, or long wait times for shared resources. At Yotta Labs, we believe infrastructure should enable discovery — not stand in its way. Today, we’re excited to announce the launch of the Yotta Labs Academic Research Support Program, an initiative designed to provide researchers with access to modern, production-grade AI infrastructure, backed by dedicated support and flexible pricing.
Research

Nov 12, 2025
NeuronMM: High-Performance Matrix Multiplication for LLM Inference on AWS Trainium
Enabling high-performance of AI workloads on heterogeneous hardware is one of the major missions at Yotta Labs. Yotta Labs has explored various AI accelerators (such as NVIDIA GPU, AMD GPU, and AWS Trainium) to optimize performance and reduce production costs. Recently, our chief scientist Dong Li, leading a team of researchers, made significant breakthroughs in building high-performance matrix multiplication (matmul) for LLM inference on Trainium. Evaluating with nine datasets and four recent LLMs, we show that NeuronMM largely outperforms the state-of–the-art matmul implemented by AWS on Trainium: at the level of matmul kernel, NeuronMM achieves an average 1.35× speedup (up to 2.22×), which translates to an average 1.66× speedup (up to 2.49×) for end-to-end LLM inference. The code is released at https://github.com/PASAUCMerced/NeuronMM.
Research

Oct 23, 2025
Optimizing Distributed Inference Kernels for AMD DEVELOPER CHALLENGE 2025: All-to-All, GEMM-ReduceScatter, and AllGather-GEMM
This technical report presents our optimization work for the AMD Developer Challenge 2025: Distributed Inference Kernels, where we develop high-performance implementations of three critical distributed GPU kernels for single-node 8× AMD MI300X configurations. We optimize All-to-All communication for Mixture-of-Experts (MoE) models, GEMM-ReduceScatter, and AllGather-GEMM kernels through fine-grained per-token synchronization, kernel fusion techniques, and hardware-aware optimizations that leverage MI300X's 8 XCD architecture. These optimizations demonstrate significant performance improvements through communication-computation overlap, reduced memory allocations, and ROCm-specific tuning, providing practical insights for developers working with distributed kernels on AMD GPUs.
Research

Oct 13, 2025
Performance Optimization for Reinforcement Learning on AMD GPUs
This blog presents our performance optimization and parameter tuning methodology for Reinforcement Learning (RL) workloads using the Verl framework on AMD’s MI300X GPU platform. By capitalizing on the MI300X’s 192GB of unified memory per GPU, we test and optimize the parallelism strategy to minimize the inter-GPU communication on the three phases in GPRO; we also explore the performance with various parallelisms and reveal the nontrivial relationship between the parallelism degree and performance.
News

Sep 23, 2025
NSF SBIR | Decentralized Artificial Intelligence (AI) Computing Operating System for Accessible and Cost-Effective AI
Yotta Labs Awarded Competitive Grant from the U.S. National Science Foundation to Advance Decentralized AI for Accessible and Cost-Effective Computing
Inference

Oct 07, 2026
Mistral Large 4 Hardware Requirements: GPU, Memory, and What to Plan For (2026)
Mistral Large 4 is a 1.05 trillion-parameter open-weight model with weights due by the end of October. The memory math at each precision, which GPU nodes fit it, and what to have ready.
GPU Pods
Distributed Inference
Inference

Oct 07, 2026
GPT-6 Astra Price: ChatGPT Plans ($100, $200 Pro) and API Cost (2026)
GPT-6 Astra is included in ChatGPT at no extra charge as GPT-6 Pro on the $100 and $200 Pro plans, Business, and Enterprise, with a message cap. On the API it costs $10 in and $50 out per million tokens, 5x GPT-6 Sol. Every plan, every tier, and the math.
Cost Optimization
Inference

Oct 06, 2026
What Is SGLang? Architecture, Performance, and When to Use It Over vLLM (2026)
SGLang is the inference engine behind some of the largest LLM deployments in production. This guide explains how RadixAttention works, where SGLang beats vLLM, and when each engine fits your workload.
SGLang
vLLM
Inference

Oct 06, 2026
Qwen 4: Release Date, What's Confirmed, and How to Prepare (2026)
Qwen 4 hasn't launched, but Alibaba has already shown its architecture. What's confirmed, what's rumor, the likely window, and how to be ready.
Cost Optimization
Distributed Inference
Inference

Oct 05, 2026
Gemma 4 Hardware Requirements: GPU and Memory for 31B, 26B, 12B, and E4B (2026)
Google’s Gemma 4 comes in five sizes under Apache 2.0, from a phone model to a 31B that fits one H100. What each size needs in VRAM, which cards work, the Ollama and vLLM commands, and where it stands against other 2026 open models.
Cost Optimization
Inference

Oct 03, 2026
GPT-OSS 120B Hardware Requirements: GPU, Memory, and How to Run It (2026)
OpenAI's open-weight 120B fits one 80 GB GPU, and the 20B runs in 16 GB. The exact memory math, which cards work, the vLLM and Ollama commands, and where it stands against 2026 open models.
vLLM
Cost Optimization
Inference

Oct 02, 2026
Qwen 3.8 Flash-Next vs Qwen 3.8 27B: Which Open Qwen to Run (2026)
Alibaba's two open Qwen 3.8 models point in opposite directions: one fits a single GPU, the other previews Qwen 4. Hardware, license, benchmarks, and which to pick.
Cost Optimization
Inference

Oct 01, 2026
DeepSeek V4: Models, Specs, Pricing, and What's Current (2026)
DeepSeek V4 is fully shipped: V4-Pro-0813 went GA on August 13 and its MIT weights are now on Hugging Face. Specs, new pricing, and how to access.
Cost Optimization
Distributed Inference
Inference

Sep 29, 2026
Qwen 4 Max vs Plus vs Flash vs 27B: The Four Tiers Explained (2026)
Alibaba named four Qwen 4 tiers on stage at Apsara and confirmed nothing else. Here's what each tier is for, what the Qwen 3.8 lineup tells you about it, and what to have ready.
Cost Optimization
Deep dives into multi-silicon AI optimization, infrastructure architecture, and the science behind Yotta's performance breakthroughs.