AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Thinking Machines has introduced Inkling, a large-scale, open multimodal AI model on Hugging Face. It processes text, images, and audio but requires substantial hardware, and its benchmarks and licensing remain unverified.

Thinking Machines has made available Inkling, a 975-billion-parameter multimodal AI model, on Hugging Face. This development is discussed in What Thinking Machines’ Inkling Signals About AI’s Evolution. This release marks a significant step in open access to large-scale models capable of processing text, images, and audio within a single framework, despite high hardware demands.

Inkling is described as a decoder-only Mixture-of-Experts model with 975 billion total parameters and 41 billion active during processing, trained on 45 trillion tokens across multiple modalities. The model supports a one-million-token context window, enabling extensive reasoning across combined data types, as detailed in the original analysis. Its architecture employs 256 experts, using a combination of global and sliding-window attention, along with specialized modules for images and audio.

Hugging Face reports that running the full model requires approximately 2 terabytes of VRAM for BF16 precision or about 600 GB for NVFP4, making it impractical for most consumer hardware. The model is accessible via supported inference engines, including Transformers, SGLang, vLLM, and llama.cpp, and can also be accessed through hosted inference services. However, details about licensing, training data transparency, and benchmark results are not yet available, highlighting the ongoing need for transparency in large-scale AI models.

At a glance
announcementWhen: announced July 2026
The developmentThinking Machines has released Inkling on Hugging Face, offering a massive multimodal model with significant hardware requirements and limited independent evaluation.

Implications of Inkling’s Open Multimodal Release

The launch of Inkling introduces a highly scalable, open-access multimodal model capable of reasoning across text, images, and audio, which could accelerate research and applications in scientific, media, and enterprise domains. However, its substantial hardware requirements limit direct deployment to organizations with advanced infrastructure, emphasizing reliance on hosted solutions or model quantization.

Amazon

high VRAM graphics card for AI training

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background on Large-Scale Multimodal AI Models

Large language models and multimodal AI systems have rapidly advanced, with models like GPT-4 and PaLM-E demonstrating multimodal capabilities. The release of Inkling by Thinking Machines follows a trend toward open, high-parameter models, although most comparable models remain proprietary or require significant resources. Prior efforts have shown that scaling up parameters enhances performance but also raises accessibility and safety concerns.

“”This model is huge.””

— Hugging Face

Amazon

large-scale AI inference server

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unverified Aspects of Inkling’s Performance and Licensing

There are no independent benchmark results, safety evaluations, or detailed licensing terms available at this stage. The actual performance of Inkling on real-world multimodal tasks, especially video processing, remains untested. It is also unclear whether the model’s open status includes full training data access or usage restrictions.

Amazon

multimodal AI model hardware requirements

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for Testing and Evaluation of Inkling

Developers and research teams are expected to begin testing Inkling through supported inference engines, focusing on latency, accuracy, and resource consumption. Independent evaluations, safety assessments, and domain-specific fine-tuning are anticipated to clarify its capabilities and limitations. Further disclosures on licensing and benchmark results are likely in upcoming updates from Thinking Machines or Hugging Face.

Local LLM Inference Optimization: A Comprehensive Guide to Quantization, Hardware Acceleration, and Efficient Private AI Deployment

Local LLM Inference Optimization: A Comprehensive Guide to Quantization, Hardware Acceleration, and Efficient Private AI Deployment

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is Inkling?

Inkling is a large, open multimodal AI model from Thinking Machines that processes text, images, and audio, with 975 billion parameters and a one-million-token context window.

Can I run Inkling on my personal computer?

Likely not. The model requires approximately 2 TB of VRAM for full deployment, making it impractical for typical consumer hardware. Access is expected primarily through hosted inference services or specialized hardware.

Does Inkling support video processing?

While the architecture allows for image inputs with a temporal dimension, native video performance has not been evaluated, and no specific video capabilities are confirmed yet.

What are the licensing terms for Inkling?

The release is described as open, but detailed licensing, usage restrictions, and training data transparency have not been disclosed.

How does Inkling compare to other multimodal models?

Independent benchmark results and safety evaluations are not yet available, so comparisons are premature. Its scale and multimodal capacity are notable, but real-world performance remains to be validated.

Source: ThorstenMeyerAI.com

You May Also Like

Fair-value appraisals for used GPUs and AI hardware

A new manual valuation system aims to establish fair market values for used data-center GPUs and AI hardware, addressing pricing disputes in the secondary market.

Writer Ian Bogost says ‘The Small Stuff’ can help us reclaim our lives from dematerialization

Ian Bogost’s new book argues that focusing on everyday sensory experiences can help us reconnect with life amid technological dematerialization.

Boost Your Academic Success With These 13 AI Productivity Apps

Discover 13 AI tools designed to enhance student productivity, from note-taking to focus, and learn how to choose the best for your needs.

Mistral’s Robostral Navigate: A State Of The Art Robotics Navigation Model

Mistral introduces Robostral Navigate, a cutting-edge robotics navigation system designed to enhance autonomous operations across industries.