Artificial intelligence

AWS Expands AI Options and Launches New Tools for Monitoring, Inference, and Agents

AWS presented a set of updates including the availability of GPT-6 Sol, GPT-6 Luna, and Claude Opus 5.5 through Amazon Bedrock, along with tools for monitoring agents, routing model inference, and building event-driven applications and software agents.

2026-09-28
4 min read
2 views
certi.news Editorial Team
AWS Expands AI Options and Launches New Tools for Monitoring, Inference, and Agents

During the week ending September 28, 2026, AWS added a set of artificial intelligence services and infrastructure tools to its ecosystem, in a direction focused on choosing the right model for each task based on intelligence, cost, and response time, rather than using the largest available model in every case.

New models through Amazon Bedrock

Amazon Bedrock made OpenAI’s GPT-6 Sol and GPT-6 Luna models available. GPT-6 Sol targets development work and repetitive, complex operations, while GPT-6 Luna focuses on specific, repeatable tasks when run at large volumes. AWS says the two models come at prices clearly lower than those of their predecessors, GPT-5.6.

Anthropic’s Claude Opus 5.5 also became available through Bedrock, as the first model in the Claude 5.5 family. It is expected to perform more tasks using fewer tokens compared with Claude Opus 5, with tuning for agent-based programming and long-running tasks.

Monitoring applications and agents in one workspace

AWS launched Amazon CloudWatch Omni to monitor applications and artificial intelligence agents within a single collaborative experience. The service is based on OpenTelemetry, allowing current telemetry data to be displayed without reconfiguration. It also provides access through a single address with unified enterprise login, without requiring direct access permissions to the console.

CloudWatch Omni automatically identifies services and maps their dependencies. It also integrates AWS DevOps Agent into investigation sessions to correlate signals and track the root causes of failures.

Practical changes in inference and event-driven architecture

AWS also introduced Amazon SageMaker HyperPod Inference Gateway, a native Kubernetes routing layer deployed as a single managed add-on for Amazon EKS, without application modifications. It routes requests according to real-time signals such as KV cache memory usage, queue length, prefix cache hit success, and expected response time, instead of traditional round-robin balancing. AWS says this reduced time to first token by up to 82% in scenarios involving diverse hardware and intermittent workloads. The layer supports OpenAI-compatible model servers, including vLLM and SGLang.

In Amazon EventBridge, it became possible to create a shared, centralized event bus across enterprise accounts through AWS RAM, with options for event ordering, content-based deduplication, and synchronous invocation of targets such as AWS Lambda. The new Subscriber resource also combines filtering policies, targets, and retry settings, while an ingress/egress pricing model replaces accumulated routing charges between accounts in multi-bus architectures. Existing buses continue to operate as “classic” buses.

What changes in practice?

These announcements combine an expanded model list with improvements to the operational layers surrounding those models. The practical value is not limited to the availability of additional models; model selection is now more closely tied to the cost and duration of a task, while CloudWatch Omni and Inference Gateway target two immediate operational problems: understanding the behavior of systems and agents, and reducing inference latency under heterogeneous workloads.

AWS also launched skills for artificial intelligence agents in Amazon SES and AWS End User Messaging through AWS MCP Server, to guide coding agents in tasks such as verifying sending identity, sending production email, and building an RCS agent with cards and buttons. These skills work with Claude Code, Codex, Cursor, and Kiro.

The Strands Agents team announced Strands harness, a general-purpose agent framework that can run locally or be deployed anywhere, and is licensed under Apache 2.0. It supports models from Amazon Bedrock, Anthropic, OpenAI, Google, and Ollama, and includes settings for managing context caching, truncating tool results, and retaining memory between runs. According to the team, it costs approximately 28% less than comparable frameworks on the same models while maintaining accuracy; this remains a claim made by the team and requires independent verification.

News source
AWS News Blog
Open original source ↗
c
Author

certi.news Editorial Team

In the same category

You may also like

View all news