Official Amazon announcements — model releases, product launches and research, each one summarized with every source covering it, by AIBriefs. RSS
How-To·Developers·1 source
AWS argues per-million-token pricing is the wrong comparison metric for production generative AI workloads, which buy outcomes like resolved support tickets rather than tokens. The post is a technical how-to for picking among OpenAI models available on Amazon Bedrock.
How-To·Developers·1 source
The technique splits LLM prompts into a fixed context portion (instructions, reference documents, conversation history) and a variable user-input portion, routing requests so shared prefixes are reused to cut latency.
How-To·Developers·1 source
AWS blog post explains how model caching on SageMaker HyperPod shrinks the gap between pod request and serving traffic, which is dominated by two sequential downloads: the inference server container image from Amazon ECR and the model weights.
Launch·Developers·1 source
TwelveLabs Marengo Embed 3.0 is now generally available as an embedding model in Amazon Bedrock Knowledge Bases, enabling video and image search by meaning. AWS says media, sports analytics, education, security, and retail teams need to find specific moments in video, which remains largely unsearchable today.
How-To·AI Agents·2 sources
Four AWS posts cover Amazon Quick Automate, a multi-agent automation capability in Amazon Quick Suite, including an end-to-end RFI questionnaire workflow and Outlook email automation. One guide cites a Gartner study finding AI saves sellers nearly five hours a week.
How-To·Developers·1 source
Configurable, instruction-driven detector runs on any LLM managed on Amazon Bedrock, evaluated on five public PII corpora across nine LLM-based detectors including OpenAI PrivacyFilter.
How-To·AI Agents·1 source
The Agent Evaluation Metric (AEM) scores agent quality at the turn level, targeting failures single-turn evaluation misses, where one early mistake corrupts later turns. AWS applies its first dimension, correctness, in the post.
How-To·Developers·1 source
AvioBook, a Thales Group company, uses Amazon Bedrock AgentCore to turn operational data into airline turnaround insights, where a few minutes of delay at one gate cascades through schedules.
How-To·Developers·1 source
Alibaba's Qwen team released Qwen3.8-2.4T-A95B on August 12, 2026 — the first Qwen-Max-class model shipped as open weights, with 2.4T total parameters and 95B activated per token. AWS published a guide for serving it on SageMaker HyperPod via vLLM.
Analysis·Business·1 source
Heurist Finance's flagship product brings institutional-style workflows into one chat experience, gathering market data, reading filings and news, running deep research, and building investment strategies for retail investors.
How-To·Developers·1 source
TorchServe is no longer actively maintained, with no planned updates, bug fixes, or security patches. AWS introduces Ray Serve Deep Learning Containers to support and simplify migration of TorchServe workloads.
Launch·AI Models·2 sources
OpenAI's GPT-6 Astra is now available on Amazon Bedrock, running on the Bedrock inference engine built for high performance, security, and scale. AWS says organizations are already running agents that write code, analyze data, and automate complex workflows on it.
Analysis·AI Models·1 source
AWS blog details how Pathway's brain-inspired architecture is developed on SageMaker HyperPod, focusing on distributed training with PyTorch. The approach contrasts with chain-of-thought by externalizing reasoning differently.
Launch·Developers·1 source
Amazon SageMaker Feature Store now supports feature-level writes via UpdateRecord, enabling updates to individual features without rewriting entire records. The feature is available in the fully managed ML feature repository.
How-To·Developers·2 sources
AWS published a two-part blog series on governing models with MLflow and Amazon SageMaker AI Model Registry sync. Part 1 covers automating model registration between MLflow and the registry; Part 2 extends governance across accounts.
Analysis·Developers·1 source
AWS compares GPU instances for small LLM inference, showing a generation jump can slash latency, increase throughput, and reduce cost-per-token. The post provides real-world benchmarks on SageMaker AI.
Analysis·Developers·1 source
HPE Zerto built an agentic troubleshooting system using Amazon Bedrock to assess health, investigate issues, and act on problems across hybrid and multi-cloud infrastructures. The post was co-written by AWS and the HPE Zerto team.
Analysis·Developers·1 source
DiDi's International Business Group built an intelligent contact center QA system on Amazon Bedrock, covering Spanish and Portuguese across ride-hailing, food delivery, and financial services.
How-To·Developers·1 source
AWS blog post demonstrates deploying a multimodal WhatsApp ordering assistant using Amazon Bedrock AgentCore and Amazon Nova 2, targeting quick-service restaurants that manage ordering across multiple channels.
How-To·Developers·4 sources
AWS AI Blog released four technical guides on Amazon Bedrock AgentCore, covering memory lifecycle policies, AI-driven development, migrating agentic workloads, and auto-generating architecture diagrams. The posts target advanced users and include practical how-to content.
How-To·Developers·1 source
AWS blog details a continuous pipeline for Physical AI systems (robots, AVs) using NVIDIA Cosmos 3 on SageMaker HyperPod, covering synthetic data generation and post-training of perception and policy models.
Launch·Developers·1 source
AWS introduces InstantStart for SageMaker HyperPod, enabling agent-driven operations that automate the chain of dependent tasks like network setup and accelerator capacity attachment. It targets foundation model workloads.
How-To·Developers·1 source
AWS blog details customizing Bedrock knowledge bases for large, complex documents using Amazon Textract. It addresses parsing multi-page utility bills with inconsistent formats and dense tables.
Analysis·Developers·1 source
Intuit developed an agentic disaster recovery assistant using Amazon Bedrock to coordinate failover across thousands of microservices spanning multiple AWS Regions. The solution addresses the operational challenge of reliable failover at scale.
How-To·Developers·1 source
AWS published a walkthrough for connecting Outlook to Amazon Quick to automate email workflows. It cites a Gartner study finding AI saves sellers nearly five hours a week.
How-To·Developers·1 source
AWS published a guide for deploying OpenAI ChatGPT Codex with LiteLLM on Amazon ECS and Amazon Bedrock, aimed at providing centralized enterprise controls for generative AI coding agents. The setup helps developers understand repositories, write code, run tests, and complete multi-step engineering tasks.
Launch·AI Models·1 source
Amazon Bedrock now offers OpenAI GPT-5.6 Sol, Terra, and Luna with global cross-Region inference from Asia Pacific (Sydney) and Asia Pacific (Melbourne) AWS Regions in Australia.
How-To·Developers·1 source
AWS AI Blog outlines using generative AI to handle rising ticket volumes, meet SLAs, and keep documentation current without proportional headcount increases. The post is a technical guide for modernizing support operations on AWS.
Analysis·Developers·1 source
AWS describes a solution using Amazon Bedrock to detect dashboard content failures, such as blank charts, that infrastructure monitors miss. The approach addresses BI dashboards where pipelines run but content fails to render.
Launch·Education·1 source
Trinity is a conversational AI solution built on Amazon Bedrock AgentCore that helps students with disabilities take ownership of postsecondary planning. It was co-authored with University Startups.
Launch·AI Models·15 sources
Fable 5.1 scores 55.8% on Terminal-Bench 4.0 in Claude Code, ahead of Fable 5 (42%) and Opus 5 (52.3%). Priced the same as Fable 5, with API cache reads cut 75% to $0.25/MTok.
Analysis·Developers·1 source
Atos upskilled 400 engineers in agentic AI, moving from theory to delivery. The program used AWS services like Amazon Bedrock and SageMaker to build real-world capability.
How-To·Developers·1 source
AWS post walks through permission models for Amazon Quick Agents, Flows, and Spaces, noting pilots that work for ten users often break when five departments are added. Agents can return data outside their intended scope.
How-To·Developers·1 source
AWS blog shows how to connect an AgentCore Runtime hosted MCP server to Amazon Quick, enabling standardized, secure access to files, databases, and APIs for AI agents.
Analysis·Developers·1 source
AWS received the highest score in the Strategy category in The Forrester Wave: AI Infrastructure Solutions, Q4 2025, which evaluated 13 providers. The recognition reflects AWS's commitment to flexible AI infrastructure.
Launch·Developers·8 sources
AWS Agent Registry manages agents, tools, and skills at scale, addressing discovery and governance challenges. The Agentic Resource Discovery (ARD) open specification enables cross-environment discovery for agents.
Launch·Developers·1 source
Amazon SageMaker Feature Store now supports batch write and record discovery, enabling efficient ingestion and retrieval of features for ML models. The feature is available in the AWS Management Console and via the SDK.
Analysis·AI Models·1 source
Decathlon, a sporting goods retailer with 400 million users, uses Chronos-2 for demand forecasting at scale. The AWS blog post details the implementation and benefits.
Analysis·Developers·1 source
Salesforce made Agentforce highly available across multiple Availability Zones using Amazon SageMaker AI Inference Components, which cut GPU costs but required custom placement to guarantee Multi-AZ redundancy.
Launch·AI Models·1 source
Amazon Bedrock now supports OpenAI GPT-5.6 models, Terra and Luna, in India with India geographic cross-Region inference. This enables local data processing for financial services, healthcare, and public sector.
Launch·Developers·1 source
Deepgram's self-hosted speech AI now exposes billing, capacity, and cost metrics through Amazon SageMaker AI Enhanced Metrics, addressing the observability trade-off of vendor-locked containers. The integration enables capacity planning and cost management directly from SageMaker.
How-To·Developers·1 source
AWS, NVIDIA, and Heidi detail how NVIDIA MPS on Amazon EC2 reduces automatic speech recognition (ASR) inference costs by 75% while meeting strict latency requirements. The post targets low GPU utilization per request.
Launch·Developers·9 sources
AWS announced new Amazon Bedrock AgentCore capabilities: AgentCore Evaluations for testing any agent framework, cross-account knowledge base connections, a Gateway for governing tool access, natural-language policy authoring, web search domain/date filters, and async invocation patterns. McKinsey reports ~80% of organizations have encountered risky AI agent behavior.
Analysis·Health·1 source
Natera, a global diagnostics company, uses Amazon Bedrock AgentCore to power intelligent appointment scheduling, allowing phlebotomists to visit oncology patients at home. The solution improves convenience for patients managing treatment.
How-To·Developers·1 source
AWS updated its 2021 script mode guide for SageMaker AI SDK v3, letting users run custom training and inference code on managed framework containers without building Docker images. The post covers setup and usage for bring-your-own-model workflows.
How-To·AI Models·2 sources
AWS AI Blog released a two-part series on data preparation for supervised fine-tuning (SFT). Part 1 covers formatting and quality; Part 2 covers advanced strategies like data selection and generating high-quality examples.
Launch·Developers·1 source
AWS announced new Ray capabilities on SageMaker HyperPod, integrating Ray with the purpose-built infrastructure for foundation model training and serving. Ray is an open-source framework for scaling distributed Python workloads.
How-To·Developers·1 source
AWS AI Blog details building an AI-powered knowledge management system using Amazon Bedrock Knowledge Bases to preserve institutional knowledge. The post addresses the challenge of 'tribal knowledge' loss when employees leave.
How-To·Developers·1 source
AWS AI Blog published a guide on building a restaurant telephony AI host using Amazon Connect, addressing phone-order handling and reducing staff burden. The post includes technical implementation details for developers.
How-To·Developers·1 source
AWS blog details AI-powered metadata correction and harmonization using Amazon Bedrock. It addresses the bottleneck of standardizing raw data as data generation accelerates.
Launch·Developers·1 source
ADOP on AWS aims to cut data engineering timelines from weeks to hours by automating ETL, quality checks, semantic models, and compliance validation. It is designed for teams standing up new data sources.
How-To·Developers·1 source
AWS introduces query-aware compression on Amazon Bedrock to cut input tokens sent to foundation models in RAG workloads, lowering costs at scale. The technique compresses context based on the query before it reaches the model.
Analysis·Developers·1 source
Panasonic Avionics uses agentic AI on AWS to accelerate diagnostics of in-flight entertainment and connectivity (IFEC) systems, which serve hundreds of airlines and billions of passengers annually.
Launch·AI Models·1 source
Amazon Bedrock now offers OpenAI GPT-5.6 models in more than 25 AWS Regions with cross-Region inference. Three variants—Sol, Terra, and Luna—support the feature, each tuned for a different balance of performance.
How-To·Developers·3 sources
Three-part AWS tutorial covers setting up a Snowflake environment, connecting it to Amazon SageMaker Canvas for data preparation and model building, and visualizing insights in Amazon QuickSight.
Analysis·Developers·1 source
AWS's blog post (Part 2 of a series) details architectural patterns for operating multi-agent systems at scale, emphasizing flexibility and avoiding vendor lock-in. It targets ML teams using Amazon Bedrock, Bedrock AgentCore, and SageMaker AI.
Analysis·Developers·1 source
AWS blog outlines how Amazon Bedrock AgentCore can scale cloud migrations, addressing bottlenecks like weeks-long discovery and manual infrastructure code. It positions agentic AI to automate migration workflows.
Analysis·Developers·1 source
AWS highlights vector search as the retrieval layer for agentic AI, enabling accurate, contextual, and grounded agents. The post covers vector solutions across Amazon Aurora, DynamoDB, ElastiCache, Neptune Analytics, OpenSearch Service, and S3.
How-To·Health·1 source
AWS shows how to use Amazon Bedrock to build intelligent security for healthcare FHIR APIs, balancing open patient data access with strict data protection. The approach reduces manual maintenance of static security rules and compliance gaps.
How-To·Developers·1 source
The walkthrough covers classifying, extracting, and validating mortgage documents such as W-2s, bank statements, and driver's licenses. It targets intermediate users automating high-volume loan document processing.