Blog_dumb

Manage agents, tools and skills at scale with AWS Agent Registry

Manage agents, tools and skills at scale with AWS Agent Registry

Most organizations scaling their use of agents and tools hit the same challenges. Teams build in isolation, with no shared record of what exists, who owns it, or whether it’s been reviewed. The problem has moved from building agents and tools to discovering and governing them. AWS Agent Registry is purpose-built to solve this. Now […]

Manage agents, tools and skills at scale with AWS Agent Registry Read More »

Build observable enterprise agentic retrieval using Managed Amazon Bedrock Knowledge Base with AWS CloudFormation

Build observable enterprise agentic retrieval using Managed Amazon Bedrock Knowledge Base with AWS CloudFormation

Teams that add Retrieval Augmented Generation (RAG) to a foundation model usually start with a single retrieval step against a single knowledge base. That works until the questions get harder, when the answer spans several sources, or the system has to decide which source to consult before it can respond. Enterprise agentic retrieval solves that:

Build observable enterprise agentic retrieval using Managed Amazon Bedrock Knowledge Base with AWS CloudFormation Read More »

Build multi-tenant agentic chat applications on enterprise data with Amazon Bedrock Managed Knowledge Base

Build multi-tenant agentic chat applications on enterprise data with Amazon Bedrock Managed Knowledge Base

Multi-tenant agentic chat assistants have become a frequent request for large-scale customers, and document chat sits at the top of the list. A user uploads a contract, a report, or a product manual, and then researches or asks questions about it immediately or in the future. The conversational interface is straightforward to build, but the

Build multi-tenant agentic chat applications on enterprise data with Amazon Bedrock Managed Knowledge Base Read More »

Batch write and discover records in Amazon SageMaker Feature Store

Batch write and discover records in Amazon SageMaker Feature Store

Amazon SageMaker Feature Store is a fully managed, purpose-built repository to store, share, and manage features for machine learning (ML) models. It provides low-latency online serving for real-time inference, an offline store for historical retention and training feature data, and supports both streaming and batch ingestion patterns. As ML platforms mature, two operational gaps surface

Batch write and discover records in Amazon SageMaker Feature Store Read More »

How Decathlon runs demand forecasting at scale with Chronos-2

How Decathlon runs demand forecasting at scale with Chronos-2

This post is co-written with Vianney Bruned, Filippo Giruzzi, Belkiss Saidi, and Carlos Ramirez from Decathlon. Decathlon is one of the world’s largest sporting goods retailers, with more than 100,000 teammates and 400 million users worldwide. The company relies on accurate demand forecasting at scale to support the availability of the appropriate products in each

How Decathlon runs demand forecasting at scale with Chronos-2 Read More »

Spreading the load: How Salesforce met Multi-AZ HA with SageMaker Inference Components

Spreading the load: How Salesforce met Multi-AZ HA with SageMaker Inference Components

When Salesforce set out to make Agentforce (Salesforce’s AI foundation for agents) highly available (HA) across multiple Availability Zones (AZs), the team faced a gap. Amazon SageMaker AI Inference Components (ICs) could cut GPU costs, but their default placement didn’t guarantee the Multi-AZ resilience Salesforce’s compliance bar required. For Salesforce, the ICs delivered an 8x

Spreading the load: How Salesforce met Multi-AZ HA with SageMaker Inference Components Read More »

Build agentic creative workflows with Amazon Quick and fal

Build agentic creative workflows with Amazon Quick and fal

Creative teams face growing demand for more assets, formats, and revisions, while their scripts, references, models, and outputs often remain fragmented across tools. Creators must repeatedly transfer context and assemble results manually. With 78% of creative leaders saying demand exceeds their teams’ capacity, faster generation alone does not solve the underlying workflow problem. To address

Build agentic creative workflows with Amazon Quick and fal Read More »

Introducing India cross-Region inference for OpenAI GPT-5.6 models on Amazon Bedrock

Introducing India cross-Region inference for OpenAI GPT-5.6 models on Amazon Bedrock

Amazon Bedrock now supports the OpenAI GPT-5.6 models, Terra and Luna, in India, with India geographic cross-Region inference. If you have local data processing requirements in India, including in financial services, healthcare, and the public sector, you can now use these OpenAI models at scale. Amazon Bedrock processes inference requests and data within India. Both

Introducing India cross-Region inference for OpenAI GPT-5.6 models on Amazon Bedrock Read More »

Deepgram deepens Amazon SageMaker AI observability with Enhanced Metrics

Deepgram deepens Amazon SageMaker AI observability with Enhanced Metrics

Self-hosted speech AI has historically carried an observability trade-off. The service can tell you an endpoint is up and how many requests it served. The questions that actually drive capacity planning and cost management stay locked inside the vendor’s container: what you are billed for, which features your traffic uses, and what the inference engine

Deepgram deepens Amazon SageMaker AI observability with Enhanced Metrics Read More »

Reduce ASR inference costs by 75% with NVIDIA MPS on Amazon EC2

Reduce ASR inference costs by 75% with NVIDIA MPS on Amazon EC2

This post is a collaboration between AWS, NVIDIA and Heidi. Reducing automatic speech recognition (ASR) inference costs on Amazon Elastic Compute Cloud (Amazon EC2) becomes critical when GPU utilization per request is low but latency requirements are strict. A single ASR inference request typically uses only 15–20 percent of a GPU’s compute capacity, yet the default

Reduce ASR inference costs by 75% with NVIDIA MPS on Amazon EC2 Read More »

Scroll to Top