Blog_dumb

Amazon SageMaker Inference: 2026 year-to-date launches in review

Amazon SageMaker Inference: 2026 year-to-date launches in review

Generative AI inference is uniquely hard: models are tens to hundreds of gigabytes, latency requirements are measured in tokens per second, cold starts can span multiple minutes as containers and weights transfer, GPU capacity is constrained, and traditional monitoring tools expose none of the token-level signals that matter in production. Amazon SageMaker AI offers customers […]

Amazon SageMaker Inference: 2026 year-to-date launches in review Read More »

Migrating multi-model AI agents to Amazon Bedrock AgentCore runtime

Migrating multi-model AI agents to Amazon Bedrock AgentCore runtime

Organizations building multi-model agentic AI applications face growing infrastructure complexity. Managing container orchestration, scaling policies, identity, and observability for multiple model types adds operational overhead. Teams often spend more time on infrastructure than on agent logic development. Developers running agentic frameworks on self-managed infrastructure such as Amazon Elastic Container Service (Amazon ECS) with AWS Fargate

Migrating multi-model AI agents to Amazon Bedrock AgentCore runtime Read More »

The new AgentCore runtime: Elastic, optimized, and consistently fast starts

The new AgentCore runtime: Elastic, optimized, and consistently fast starts

Agents are no longer experiments. They process claims, write and review code, coordinate across systems, and run for hours without supervision. As agents take on more complex, longer-running work, the infrastructure underneath them must evolve just as fast. We built Amazon Bedrock AgentCore to help developers build, connect, and optimize agents securely at scale. AgentCore

The new AgentCore runtime: Elastic, optimized, and consistently fast starts Read More »

Deploy Hugging Face models on Amazon SageMaker AI with coding agents

Deploy Hugging Face models on Amazon SageMaker AI with coding agents

Deploying a Hugging Face model to production means making a dozen decisions: choosing the right serving container for the model’s architecture, confirming the current image tag for your AWS Region, and matching an instance type to the model’s memory footprint. Beyond infrastructure, you must wire autoscaling so you don’t burn GPU hours on an idle

Deploy Hugging Face models on Amazon SageMaker AI with coding agents Read More »

Introducing Amazon SageMaker HyperPod Inference Gateway

Introducing Amazon SageMaker HyperPod Inference Gateway

Eliminate GPU waste. Reduce first-token latency by up to 82%. Install one Kubernetes-native addon with zero application changes. The problem: Naive routing wastes your most expensive resource Running large language models (LLMs) at scale on GPU clusters is expensive. The default Kubernetes load balancers are making it worse. Round-robin and least-connections algorithms have no visibility

Introducing Amazon SageMaker HyperPod Inference Gateway Read More »

Reduce time-to-hire for quality candidates with AI-powered Amazon Connect Talent

Reduce time-to-hire for quality candidates with AI-powered Amazon Connect Talent

Hiring at scale in industries such as retail, logistics, hospitality, and others has its fair share of challenges. Recruiting teams are expected to fill hundreds of roles within tight timelines, often with limited capacity and with tools that weren’t designed to seamlessly work together. As a result of this, applications pile up, phone screens get

Reduce time-to-hire for quality candidates with AI-powered Amazon Connect Talent Read More »

Scroll to Top