Amazon Bedrock Kimi K3 became generally available on 18 September, giving AWS customers managed access to Moonshot AI’s open-weight model through US and global cross-region inference profiles. The integration supports text and image input, a one-million-token context window and explicit prompt caching for repeated context.
- Kimi K3 is not a new model release; the new event is managed Bedrock availability.
- Explicit prompt caching can reduce repeated processing in long coding and knowledge workflows, but only when enough context stays stable.
- Teams must test API differences, region routing, pricing and data controls rather than assuming the open weights behave identically across hosts.
Everyone else is reporting another model in a catalog; Lapaas Voice is explaining when the cache and managed boundary change deployment economics.
Amazon Bedrock Kimi K3 adds a managed path
Moonshot released Kimi K3 earlier in 2026. Bedrock changes how an enterprise can consume it: requests go through AWS’s managed runtime, access controls, encryption, logging and cross-region routing instead of a self-hosted cluster or Moonshot’s own service.
AWS documents Global and US geographic inference profiles. That distinction matters for teams with residency requirements because a global profile can route among supported commercial regions, while the US profile remains within the stated geography. Buyers should select the profile explicitly and confirm where logs, prompts and completions are processed.
Prompt caching is the practical differentiator
AWS says Kimi K3 is the first open-weight model on Bedrock with explicit prompt caching. The model card lists a 1,024-token minimum checkpoint and at least 30 minutes of retention. Cache reads are priced below ordinary input processing, while cache writes carry their own cost.
That structure favours a coding agent repeatedly reading the same repository map, or an analyst asking several questions over a stable document collection. It helps less when every prompt is short or constantly rewritten. Teams should measure cache-hit rate, write frequency and total tokens before projecting savings.
The mechanism connects with Dream-RSI agent search replay: both reduce repeated computation by reusing prior work. The difference is that a prompt cache reuses input state, while search replay reuses a trajectory. Neither removes the need to validate output quality.
Limits to include in testing
AWS lists a one-million-token context window and native vision, but its Bedrock documentation says video is not supported. It also notes differences among Responses, Chat Completions, Converse and Invoke APIs. Developers should test tool use, structured output, image ordering and cache controls on the exact API selected for production.
Claims about model scale and benchmark performance originate with Moonshot or benchmark operators and should remain attributed. Managed availability does not independently validate accuracy, security or suitability for sensitive decisions.
In short: Amazon Bedrock Kimi K3 matters less as another model listing than as a governed inference route with reusable long context. Its adoption case rests on real cache reuse, regional requirements and application-level evaluation.
Related Lapaas Voice coverage: Dream-RSI agent search replay and PeakMetrics AI Perceptions.
| Item | Verified detail |
|---|---|
| Bedrock availability | 18 September 2026 |
| Model provider | Moonshot AI |
| Context window | 1 million tokens |
| Input | Text and images; no video on Bedrock |
| Cache feature | Explicit prompt caching |
| Inference profiles | US Geo and Global cross-region |
Frequently asked questions
What is Amazon Bedrock Kimi K3?
It is Moonshot AI’s open-weight Kimi K3 model offered through AWS-managed Bedrock inference profiles.
Why does prompt caching matter?
Repeated prompt content can be reused instead of processed from scratch, which may reduce latency and input cost for stable long contexts.
Can Kimi K3 process video on Bedrock?
No. AWS documentation says video input is not supported; text and image workflows are the relevant modalities.
Sources
- AWS What’s New (2026-09-18; primary_product_announcement)
- AWS Kimi K3 model card (2026-09-18; primary_product_documentation)
- AI Stack Current (2026-09-19; independent_technical_report)
- VentureBeat (2026-07-16; independent_model_context)
Get the day’s top stories in your inbox
One concise email. No spam, unsubscribe anytime.



