Amazon Web Services (AWS) has brought OpenAI’s GPT-5.6 Terra and GPT-5.6 Luna models to India through Amazon Bedrock, giving businesses access to OpenAI’s latest models with in-country inference. The move means customer data used for model inference can be processed on AWS infrastructure in India, potentially making the models more attractive to organisations with data-residency, governance and latency requirements.

The India rollout also comes with significantly lower model pricing. OpenAI cut the price of GPT-5.6 Luna by 80% and GPT-5.6 Terra by 20% in July, making lower-cost AI inference available just as AWS expands access to the models on Bedrock. The combination of local processing, enterprise controls and lower token costs could make large-scale AI deployment more economical for Indian companies.

AWS Brings GPT-5.6 To India

AWS has made OpenAI’s GPT-5.6 Terra and GPT-5.6 Luna generally available on Amazon Bedrock in India. The models are being offered with in-country inference, meaning inference is processed on AWS infrastructure within India.

Amazon said the availability is designed to give organisations access to frontier AI while addressing requirements around data residency, low latency, security and enterprise deployment. Customers can use the models through the Bedrock environment alongside models from other AI companies.

The launch expands the partnership between AWS and OpenAI and gives Indian enterprises another route to deploy OpenAI technology without having to build a separate infrastructure stack around the models.

Key Numbers At A Glance

MetricDetails
Models launched in IndiaGPT-5.6 Terra and GPT-5.6 Luna
AWS platformAmazon Bedrock
Inference locationIndia
Maximum reported price reduction80%
Luna price reduction80%
Terra price reduction20%
GPT-5.6 Luna new input price$0.20 per 1M tokens
GPT-5.6 Luna new output price$1.20 per 1M tokens
GPT-5.6 Terra new input price$2 per 1M tokens
GPT-5.6 Terra new output price$12 per 1M tokens
India availabilityGeneral availability

GPT-5.6 Prices Fall Sharply

The biggest change for developers and businesses is the reduction in inference costs.

OpenAI announced on July 30 that GPT-5.6 Luna would become 80% cheaper, while GPT-5.6 Terra would receive a 20% price reduction. OpenAI said the reductions were made possible by improvements in how the models are built and served.

Luna was initially priced at $1 per million input tokens and $6 per million output tokens. Following the reduction, those prices fell to $0.20 and $1.20 respectively.

Terra, meanwhile, moved from $2.50 per million input tokens and $15 per million output tokens to $2 and $12.

GPT-5.6 Pricing Before And After The Cuts

ModelEarlier Input Price / 1M TokensNew Input Price / 1M TokensEarlier Output Price / 1M TokensNew Output Price / 1M TokensReduction
GPT-5.6 Luna$1.00$0.20$6.00$1.2080%
GPT-5.6 Terra$2.50$2.00$15.00$12.0020%
GPT-5.6 Sol$5.00$5.00$30.00$30.00No announced cut

The difference is particularly significant for high-volume applications. Businesses running large numbers of automated queries, customer-service interactions, document-processing tasks or AI agents can see substantial changes in operating costs when the underlying token price falls.

Why The 80% Luna Price Cut Matters

GPT-5.6 Luna is positioned as the fastest and most affordable model in the GPT-5.6 family. OpenAI says it is designed for high-volume work where organisations need strong performance at a lower cost.

The lower price could make workloads that were previously too expensive to automate more commercially viable.

For example, companies can use lower-cost models for repetitive tasks such as classification, summarisation, extraction, customer-support responses and multi-step workflows, while reserving more expensive models for complex reasoning.

Potential Enterprise Use Cases

Use CaseWhy Lower Pricing Matters
Customer supportMore AI responses can be processed within the same budget
Document processingLarge volumes of documents become cheaper to analyse
Coding assistanceDevelopers can run more AI-assisted workflows
Data extractionHigh-volume structured extraction becomes more economical
AI agentsMulti-step workflows can run at lower inference costs
Content operationsBusinesses can process larger volumes of text
Internal knowledge toolsMore employees can access AI-powered systems

This could be particularly important for Indian companies, where large user populations and cost-sensitive technology budgets can make AI inference costs a major factor in deployment decisions.

In-Country Inference Is A Major Part Of The India Launch

The pricing reduction is only one part of the announcement.

AWS is also emphasising in-country inference, with OpenAI model processing taking place on AWS infrastructure in India. This is important for businesses that need to manage where sensitive information is processed.

Financial institutions, healthcare organisations, government-related entities and large enterprises can face strict requirements concerning data handling and governance. Having model inference processed within the country can make deployment easier for organisations with such requirements.

What In-Country Inference Offers

FeaturePotential Enterprise Benefit
Processing in IndiaSupports local data-residency requirements
AWS infrastructureUses existing enterprise cloud environment
Lower latencyPotentially faster interaction for Indian workloads
Security controlsFits into existing AWS governance frameworks
Bedrock integrationEasier access alongside other foundation models
Centralised managementSimplifies model deployment and monitoring

AWS says the models are available through Amazon Bedrock’s existing controls and infrastructure, allowing customers to use OpenAI models alongside models from other providers.

How GPT-5.6 Terra And Luna Differ

The two models are aimed at somewhat different workloads.

Terra is positioned as a balanced model for everyday production work, while Luna is designed to offer faster and more affordable inference for high-volume applications. OpenAI’s broader GPT-5.6 family also includes Sol, its flagship model for more demanding workloads.

GPT-5.6 Model Lineup

ModelPositioningBest Suited ForAPI Input Price / 1M Tokens
GPT-5.6 SolFlagshipComplex reasoning, coding, research$5
GPT-5.6 TerraBalancedEveryday production workloads$2
GPT-5.6 LunaFast, low-costHigh-volume AI applications$0.20

This tiered approach allows companies to match model capability with the economic value of each task instead of using the most expensive model for every workload.

AWS Bedrock Expands OpenAI Access For Enterprises

Amazon Bedrock has increasingly become a model-access layer through which businesses can evaluate and deploy foundation models from multiple providers.

AWS says customers can use OpenAI models alongside models from Anthropic, Meta, Mistral, Cohere, Amazon and other providers through a common service with unified security, governance and cost controls.

For companies already using AWS, this can reduce the operational friction involved in introducing another AI model into their technology stack.

Instead of managing separate infrastructure and integrations for every model provider, businesses can evaluate different models through the same cloud environment and select models according to performance, price and workload requirements.

Lower AI Costs Could Accelerate Adoption In India

The combination of lower prices and local inference could be particularly important for India’s rapidly expanding AI market.

AI adoption is moving beyond experimental chatbots toward automated business processes, coding agents, customer-service systems, enterprise search and data-analysis tools. These applications can generate millions or billions of tokens, making inference economics increasingly important.

An 80% price reduction does not necessarily mean an organisation’s overall AI bill will fall by 80%. Companies may increase usage as costs decline, while application architecture, input and output token volumes, storage, networking and other cloud services also contribute to total costs.

However, the lower model price can change the economics of what businesses are willing to automate.

OpenAI’s Broader Push Toward AI Efficiency

The India launch comes after OpenAI made efficiency and price-performance a major focus for GPT-5.6.

OpenAI said GPT-5.6 was designed to produce more useful work from every token. It also said GPT-5.6 Luna and Terra were made cheaper after improvements to the way the models are served.

The company has also highlighted strong performance from the GPT-5.6 family across coding, professional knowledge work, browsing, science and other workloads. GPT-5.6 Luna is intended to extend those capabilities into applications where cost and speed are more important than maximum reasoning performance.

The broader industry is moving in the same direction, with AI companies competing increasingly on both model intelligence and cost per task.

What It Means For Indian Businesses

For Indian enterprises, the AWS rollout creates three important advantages: access to OpenAI models through an established cloud platform, local inference and significantly lower pricing for two GPT-5.6 models.

The biggest potential beneficiaries could be businesses operating large-scale AI applications where inference costs represent a significant share of technology spending.

Startups could also benefit because lower model costs reduce the amount of capital required to test and scale AI-powered products. Larger companies, meanwhile, can potentially expand AI deployment from small pilot projects to production workloads.

The competitive impact could extend beyond OpenAI. AWS customers can compare OpenAI’s models with other foundation models available through Bedrock, increasing pressure on providers to deliver stronger performance at lower costs.

The Bigger Picture

AWS’s GPT-5.6 rollout in India illustrates how the AI market is shifting from simply making increasingly capable models available to making them affordable and practical at enterprise scale. The combination of lower inference prices, cloud integration and local processing could help turn more AI experiments into production applications.

For India, the local availability of OpenAI models on AWS also strengthens the country’s position as a major enterprise AI market. As businesses become more comfortable with AI agents and automated workflows, the cost of running those systems will become as important as model quality. The 80% reduction in Luna pricing therefore matters not only as a price cut, but as a potential catalyst for broader AI usage.

Looking Ahead

The next phase of competition will likely focus on price per completed task rather than price per token alone. Businesses will increasingly compare how much useful work a model can complete for a given budget, taking into account accuracy, latency, tool use and the amount of human intervention required. AWS’s India rollout gives enterprises another option for making that calculation within an existing cloud environment.

As AI adoption expands, lower inference costs could encourage companies to deploy models across more business processes. The combination of GPT-5.6’s lower pricing and AWS’s in-country infrastructure could make high-volume AI applications more viable in India, while also intensifying competition among OpenAI, cloud providers and rival model developers over performance, security and cost.

Get the day’s top stories in your inbox

One concise email. No spam, unsubscribe anytime.