Blog_dumb

Deploying Kimi K3 on AWS

Deploying Kimi K3 on AWS

Open weight models have become powerful enough to handle complex tasks such as multi-step agentic workflows, advanced reasoning, and long-horizon coding. However, as these models grow in capability, they also grow in size and hosting multi-trillion parameter architectures requires purpose-built infrastructure, high-end GPU compute, and optimized serving frameworks. On July 27, 2026, Moonshot AI released […]

Deploying Kimi K3 on AWS Read More »

How Yahoo enhances search retargeting using Amazon Bedrock

How Yahoo enhances search retargeting using Amazon Bedrock

Connecting user search intent with relevant ad experiences across channels is a longstanding challenge in digital advertising. Advertisers need sophisticated ways to reach audiences based on their demonstrated interests and behaviors, particularly their search activity, which is one of the strongest signals of user intent. Traditional keyword expansion approaches often struggle with outdated vocabulary, limited

How Yahoo enhances search retargeting using Amazon Bedrock Read More »

Inference meta-monitoring for Amazon SageMaker AI endpoints with Amazon Quick

Inference meta-monitoring for Amazon SageMaker AI endpoints with Amazon Quick

Without the ability to track machine learning (ML) model prediction quality, organizations only realize they have issues when their customers complain or when they conduct spot checks, which jeopardizes customer trust. This post introduces inference meta-monitoring for Amazon SageMaker AI endpoints. It provides a governance layer that sits above production ML inference pipelines to continuously

Inference meta-monitoring for Amazon SageMaker AI endpoints with Amazon Quick Read More »

Introducing explicit prompt caching for OpenAI GPT-5.6 models on Amazon Bedrock

Introducing explicit prompt caching for OpenAI GPT-5.6 models on Amazon Bedrock

This post is co-written with Chris Dickens from OpenAI. OpenAI GPT-5.6 Sol, Terra, and Luna are now generally available on Amazon Bedrock. With GPT-5.6 on Amazon Bedrock, you get the newest generation of OpenAI frontier models with pay-per-token pricing, AWS security and governance controls, and usage that counts toward your existing AWS commitments. The family

Introducing explicit prompt caching for OpenAI GPT-5.6 models on Amazon Bedrock Read More »

Migrate your prompts to new models and optimize them on Amazon Bedrock

Migrate your prompts to new models and optimize them on Amazon Bedrock

The problem: Prompt engineering at scale is still painful Migrating prompts to new models on Amazon Bedrock, or optimizing them for your current model, is still one of the most manual parts of building a generative AI application. Say you have built a deployed generative AI application. It works. Your prompts are tuned, your outputs

Migrate your prompts to new models and optimize them on Amazon Bedrock Read More »

Authenticate with Private Key JWT using Amazon Bedrock AgentCore Identity

Authenticate with Private Key JWT using Amazon Bedrock AgentCore Identity

Amazon Bedrock AgentCore Identity now supports Private Key JWT client authentication for agents. With Private Key JWT client authentication, your agents can authenticate to a downstream identity provider’s token endpoint using a signed JSON Web Token (JWT) client assertion instead of a shared OAuth 2.0 client secret. You can register a public key with your

Authenticate with Private Key JWT using Amazon Bedrock AgentCore Identity Read More »

Generate Autonomous Business Insights with AI Agent and MCP Servers

Generate Autonomous Business Insights with AI Agent and MCP Servers

A Monday morning problem Sarah Chen manages 12 assembly lines and 2,000 machines. Before her 10 AM production review, she needs one answer: Which lines need attention this week? Simple question. Painful journey. She starts in the IoT dashboard. Line 4’s motor temperature is running 12°C above baseline — has been for three days. Calibration

Generate Autonomous Business Insights with AI Agent and MCP Servers Read More »

Automating customer retention workflows in Amazon Quick

Automating customer retention workflows in Amazon Quick

Automating customer retention workflows in Amazon Quick can turn a five-day churn-response cycle into one that takes minutes. Last quarter, a mid-size SaaS company lost 12% of its at-risk accounts because the retention team took five days to identify and contact dissatisfied customers. By the time someone manually reviewed CSAT spreadsheets and call transcripts, those

Automating customer retention workflows in Amazon Quick Read More »

How AgentCore Gateway supports the MCP 2026-07-28 spec

How AgentCore Gateway supports the MCP 2026-07-28 spec

Today, the Model Context Protocol (MCP) published its 2026-07-28 specification, the largest and most significant revision of the protocol since its launch. With this release MCP becomes a stateless protocol that scales on ordinary HTTP infrastructure. Alongside the transport changes, this new version introduces a governed extensions system, strengthens authorization by aligning more closely with

How AgentCore Gateway supports the MCP 2026-07-28 spec Read More »

Market surveillance agent with LangGraph and Strands on AgentCore

Market surveillance agent with LangGraph and Strands on AgentCore

As artificial intelligence applications evolve from simple chatbots to sophisticated autonomous systems, organizations face new challenges in orchestrating complex multi-agent workflows that can handle real-world production scenarios. Traditional single-agent approaches often fall short when dealing with intricate business processes that require specialized expertise, dynamic decision-making, and robust error recovery mechanisms. The financial services industry exemplifies

Market surveillance agent with LangGraph and Strands on AgentCore Read More »

Scroll to Top