Blog_dumb

Deploying quantized models on Amazon SageMaker AI with Unsloth

Deploying quantized models on Amazon SageMaker AI with Unsloth

This post was co-written with Daniel Han and Michael Han from Unsloth. Deploying large foundation models (FMs) stored at their original 16-bit floating-point precision (BF16 or FP16) is expensive. They need large GPU instances, driving up serving costs, and slowing down iteration cycles. Quantization addresses this by reducing the numerical precision of a model’s weights […]

Deploying quantized models on Amazon SageMaker AI with Unsloth Read More »

Enhancing enterprise inference on Amazon SageMaker HyperPod with data capture, Hugging Face, NVMe, and Route 53 integration

Enhancing enterprise inference on Amazon SageMaker HyperPod with data capture, Hugging Face, NVMe, and Route 53 integration

As enterprises scale their generative AI workloads, the demand for faster, more observable, and more flexible inference infrastructure continues to grow. Amazon SageMaker HyperPod is rising to meet that challenge with a set of new capabilities designed to streamline how organizations deploy and operate large models in production. Teams can now record inputs and outputs

Enhancing enterprise inference on Amazon SageMaker HyperPod with data capture, Hugging Face, NVMe, and Route 53 integration Read More »

Powering scientific discovery: BYOKG and GraphRAG for intelligent pharmaceutical research

Powering scientific discovery: BYOKG and GraphRAG for intelligent pharmaceutical research

In pharmaceutical research, scientists face a fundamental challenge: accessing and connecting the vast amount of scientific knowledge scattered across disparate systems. From published literature and internal lab notes to genomics databases, critical insights remain trapped in silos, making it difficult for researchers to form comprehensive connections and generate promising hypotheses. This fragmentation slows down the

Powering scientific discovery: BYOKG and GraphRAG for intelligent pharmaceutical research Read More »

Automatically sort and prioritize your mailboxes by using Amazon Bedrock

Automatically sort and prioritize your mailboxes by using Amazon Bedrock

AI-powered email management can transform how organizations in the public sector handle constituent communications. By implementing intelligent email routing and prioritization systems, organizations can automatically classify and direct incoming messages based on urgency and departmental relevance. This technology is particularly useful in local government settings, where councillors receive diverse communications across multiple service areas. AI

Automatically sort and prioritize your mailboxes by using Amazon Bedrock Read More »

Building and connecting a production-ready ecommerce MCP server using Amazon Bedrock AgentCore and Mistral AI Studio

Building and connecting a production-ready ecommerce MCP server using Amazon Bedrock AgentCore and Mistral AI Studio

When ecommerce teams need faster time-to-market for AI-powered customer experiences, they face weeks of custom integration work that delays launches and increases security risks. Building and connecting a production-ready AI assistant typically requires custom API code for each client, container infrastructure management, and complex authentication. Amazon Bedrock AgentCore and Mistral AI Studio streamline this process.

Building and connecting a production-ready ecommerce MCP server using Amazon Bedrock AgentCore and Mistral AI Studio Read More »

Securing Amazon Bedrock AgentCore Runtime with AWS WAF

Securing Amazon Bedrock AgentCore Runtime with AWS WAF

When you deploy generative AI agents with Amazon Bedrock AgentCore as production API endpoints, you might want to enforce web application firewall policies, rate limiting, protection against common web threats, or audit controls via AWS WAF. AWS WAF integrates with Elastic Load Balancing Application Load Balancers (ALBs), Amazon CloudFront distributions, and Amazon API Gateway REST

Securing Amazon Bedrock AgentCore Runtime with AWS WAF Read More »

Manage AI applications on Mac with Jamf’s AI Governance and Amazon Bedrock

Manage AI applications on Mac with Jamf’s AI Governance and Amazon Bedrock

As organizations expand AI adoption across their workforce, IT administrators need a scalable way to manage how AI applications are configured and used on employee devices. These applications include Claude Code, Claude Desktop, and OpenAI Codex. Users, meanwhile, can open approved applications and start working without manual setup. Jamf, trusted by more than 78,000 organizations

Manage AI applications on Mac with Jamf’s AI Governance and Amazon Bedrock Read More »

Enrich your datasets with business context: Migrating from legacy Topics to semantic datasets in Amazon Quick

Enrich your datasets with business context: Migrating from legacy Topics to semantic datasets in Amazon Quick

If you’ve been managing Amazon Quick legacy Topics alongside your datasets, you know the challenge: two assets that must stay perfectly synchronized, each with its own permissions, lineage, and versioning. Column synonyms drift. Calculated fields diverge. A rename in the dataset breaks the Legacy Topic silently. You can now use Amazon Quick to embed that

Enrich your datasets with business context: Migrating from legacy Topics to semantic datasets in Amazon Quick Read More »

Scroll to Top