Thinking Machines Lab, the artificial intelligence startup founded by former OpenAI CTO Mira Murati, has introduced Inkling-Small, a new open-weight AI model that delivers performance close to its predecessor while requiring only about one-quarter of the model size. The release reflects a growing industry trend toward developing smaller, more efficient AI models that can provide frontier-level capabilities at significantly lower computational costs.

Inkling-Small is a Mixture-of-Experts (MoE) transformer model featuring 276 billion total parameters with only 12 billion active parameters during inference. Despite being substantially smaller than the original Inkling model, the company says it matches or exceeds its predecessor on several reasoning, coding, and agentic AI benchmarks while maintaining multimodal capabilities, including support for text, images, and audio.

Inkling-Small Delivers Comparable Performance in a Smaller Package

Thinking Machines Lab designed Inkling-Small to improve deployment efficiency without sacrificing AI performance.

According to the company, the model offers:

  • Performance close to the original Inkling model.
  • Roughly one-quarter of the predecessor’s size.
  • Lower inference costs.
  • Faster deployment on enterprise hardware.
  • Open-weight availability for developers and researchers.

Model Snapshot

FeatureInklingInkling-Small
Total Parameters975 billion276 billion
Active ParametersNot disclosed12 billion
Model TypeMixture-of-Experts (MoE)Mixture-of-Experts (MoE)
Multimodal SupportYesYes
Open WeightsYesYes

Designed for Efficient AI Deployment

Inkling-Small follows the company’s earlier release, when Thinking Machines launched Inkling, its first open-weight AI model.

Unlike traditional dense language models that activate every parameter during inference, Inkling-Small uses a Mixture-of-Experts architecture, activating only a small subset of parameters for each request.

This approach provides several advantages:

  • Lower computational requirements.
  • Reduced inference latency.
  • Lower operating costs.
  • Improved scalability for enterprise deployments.
  • Better energy efficiency.

The model also supports a context window of up to one million tokens, enabling it to process very large documents, repositories, or multimodal inputs in a single session.

Multimodal AI With Voice and Vision Capabilities

Like its predecessor, Inkling-Small is designed as a multimodal foundation model.

Supported capabilities include:

  • Natural language understanding.
  • Code generation.
  • Image understanding.
  • Audio processing.
  • Agentic reasoning.
  • Adjustable reasoning effort depending on task complexity.

Thinking Machines Lab says the model performs particularly well on reasoning and AI agent workloads while maintaining competitive vision and speech capabilities.

Key Capabilities

CapabilitySupported
Text GenerationYes
CodingYes
Image UnderstandingYes
Audio UnderstandingYes
Agentic AIYes
Long Context (1M Tokens)Yes

Open-Weight Release Targets Developers

Inkling-Small is being released as an open-weight model, allowing researchers and enterprises to deploy and fine-tune it for their own applications.

The company has made the model available through:

  • Its Tinker AI platform.
  • Hugging Face model repository.
  • Multiple precision formats, including BF16 and NVFP4, to support different hardware configurations.

The release is expected to appeal to organizations seeking greater control over AI infrastructure, data privacy, and customization compared with proprietary hosted models.

Reflecting a Broader Industry Trend

The lab’s influence has extended beyond its own releases too, after Sarvam appointed Thinking Machines Lab founding member Devendra Singh Chaplot as an advisor.

The launch highlights a broader shift in the AI industry toward smaller, highly optimized models that deliver near-frontier performance with substantially lower hardware requirements.

Rather than relying solely on ever-larger foundation models, AI companies are increasingly focusing on:

  • Model compression.
  • Efficient inference.
  • Lower deployment costs.
  • Enterprise-ready open-weight models.
  • Specialized AI agents.

This approach enables businesses to deploy advanced AI workloads more economically while maintaining strong performance across reasoning and multimodal tasks.

Looking Ahead

Inkling-Small demonstrates that advances in AI are no longer driven solely by increasing model size. By delivering performance approaching its 975-billion-parameter predecessor while operating at roughly one-quarter of the footprint, Thinking Machines Lab is emphasizing efficiency, scalability, and practical deployment. The release also strengthens the company’s position in the growing market for open-weight enterprise AI models, where organizations increasingly value flexibility, lower inference costs, and control over their AI infrastructure.

Looking ahead, competition among AI developers is likely to focus not only on benchmark performance but also on efficiency, multimodal capabilities, and deployment economics. As enterprises seek models that can run on more accessible hardware while supporting complex reasoning, coding, and multimodal workloads, compact open-weight models like Inkling-Small could play an increasingly important role in accelerating enterprise AI adoption.

Frequently Asked Questions

What is Inkling-Small?

Inkling-Small is a new open-weight AI model from Thinking Machines Lab that delivers performance close to its predecessor while requiring only about one-quarter of the model size.

Who founded Thinking Machines Lab?

Thinking Machines Lab is the artificial intelligence startup founded by former OpenAI CTO Mira Murati.

Does Inkling-Small support multimodal capabilities?

Yes, according to the article, Inkling-Small includes multimodal AI with voice and vision capabilities.

Get the day’s top stories in your inbox

One concise email. No spam, unsubscribe anytime.