Skip to content
-
  • https://www.facebook.com/
  • https://twitter.com/
  • https://t.me/
  • https://www.instagram.com/
  • https://youtube.com/
AI Feed AI Feed AI Feed

AI news, tools, comparisons and practical guides

Subscribe
AI Feed AI Feed AI Feed

AI news, tools, comparisons and practical guides

  • AI News
  • AI Tools Radar
  • AI Comparisons
  • About
  • AI API Prices

Sections

  • AI Comparisons
  • AI Features
  • AI Guides
  • AI News
  • Uncategorized

Latest stories

  • AWS Details How to Deploy Interactive MCP Apps on AgentCore
  • xAI Makes Grok 4.6 Available in GitHub Copilot
  • Hugging Face and AWS Link Strands Robots to Streaming LeRobot Training
  • Anthropic Reports Claude Results in Protein Design and Chemistry
  • xAI Opens Grok Build to Every Plan on Web and Mobile
  • AI News
  • AI Tools Radar
  • AI Comparisons
  • About
  • AI API Prices
Subscribe
Close

Search

Home/AI Features/AWS Details Qwen3.8-2.4T-A95B Deployment on HyperPod
Иллюстрация к новости: AWS описала развёртывание Qwen3.8-2.4T-A95B на SageMaker HyperPod
AI FeaturesAI News

AWS Details Qwen3.8-2.4T-A95B Deployment on HyperPod

Kat
By Kat
10.09.2026 2 Min Read
◉2unique readers

AWS published a technical guide on September 9, 2026, for self-hosting Qwen3.8-2.4T-A95B on Amazon SageMaker HyperPod. Its configuration runs vLLM on one ml.p6-b300.48xlarge instance with eight NVIDIA B300 Blackwell Ultra GPUs and exposes an OpenAI-compatible endpoint for applications.

Released by Alibaba’s Qwen team on August 12, the open-weight model has 2.4 trillion total parameters while activating about 95 billion per token, according to AWS. Its 92-layer architecture combines 69 Gated DeltaNet layers with bounded recurrent state and 23 full-attention layers. The native context window is 262,144 tokens, extensible to 1,010,000, while maximum output length is 128,000 tokens.

Model size is the central hardware constraint. AWS estimates that BF16 weights alone would occupy about 4.8 TB, exceeding the capacity of a single eight-GPU node. NVFP4 W4A4 quantization reduces the weight footprint to approximately 1.2 TB. The selected instance provides 2.1 TB of aggregate HBM3e memory, with AWS estimating roughly 500–700 GB of remaining headroom for batching, longer contexts, KV cache and other runtime needs.

The vLLM setup shards the model across all eight GPUs using tensor parallelism. AWS’s serving command also enables shared-prefix caching, automatic tool selection, Qwen3 reasoning parsing and native Multi-Token Prediction speculative decoding. The MTP implementation uses draft heads included in the model weights, so it does not require a separate draft model.

Deployment is declared through an InferenceEndpointConfig resource in a HyperPod cluster orchestrated by Amazon EKS. The HyperPod Inference Operator manages weight downloads, container scheduling, GPU allocation, readiness checks, rolling updates and endpoint lifecycle. AWS says an uncached deployment, including the roughly 1.2 TB download and weight loading, should take about 15–30 minutes depending on network bandwidth; local NVMe caching can accelerate later restarts.

The ml.p6-b300.48xlarge instance is not available on demand. Customers must obtain committed capacity through an AWS Flexible Training Plan and configure the cluster’s target Availability Zone to match the plan. In the documented example, the endpoint is not exposed to the public internet, and vLLM does not require authentication by default.

Practical context: The design is therefore a self-hosted deployment rather than a conventional managed model API: it requires a reserved eight-B300 node, terabyte-scale weight storage and transfer, Kubernetes configuration, and independent endpoint access controls. AWS characterizes the node as suitable for production inference with moderate concurrency, but the post does not provide measured throughput for this exact single-node NVFP4 configuration.

Sources

  1. AWS AI Blog

Event date: 2026-09-09. Primary source date: 2026-09-09.

Follow AI Feed on Telegram

New AI stories, practical guides and tool comparisons — in one concise feed.

Open Telegram→

Tags:

Editor’s Picks
Kat
Author

Kat

Follow Me
Other Articles
Иллюстрация к новости: Hugging Face внедрила Qwen3-Embedding-0.6B в поиск Papers with Code
Previous

Hugging Face Adds Qwen Hybrid Search to Papers with Code

Иллюстрация к новости: xAI открыла доступ к Grok 4.6 через Amazon Bedrock
Next

xAI Makes Grok 4.6 Generally Available on Amazon Bedrock

Recent posts

  • AWS Details How to Deploy Interactive MCP Apps on AgentCore
  • xAI Makes Grok 4.6 Available in GitHub Copilot
  • Hugging Face and AWS Link Strands Robots to Streaming LeRobot Training
  • Anthropic Reports Claude Results in Protein Design and Chemistry
  • xAI Opens Grok Build to Every Plan on Web and Mobile

Recent comments

No comments to show.

Archives

  • September 2026
  • May 2026

Sections

  • AI Comparisons
  • AI Features
  • AI Guides
  • AI News
  • Uncategorized

    © 2026 AI Feed. All rights reserved.
    RUEN
    AboutEditorial PolicySources & methodologyCorrectionsContactPrivacyAnalytics settings
    AI Feed analytics

    Helps us understand which pages are useful. Advertising tracking is disabled.