AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

Thinking Machines has introduced Inkling, a large-scale, open multimodal AI model on Hugging Face. It processes text, images, and audio but requires substantial hardware, and its benchmarks and licensing remain unverified.

Thinking Machines has made available Inkling, a 975-billion-parameter multimodal AI model, on Hugging Face. This development is discussed in What Thinking Machines’ Inkling Signals About AI’s Evolution. This release marks a significant step in open access to large-scale models capable of processing text, images, and audio within a single framework, despite high hardware demands.

Inkling is described as a decoder-only Mixture-of-Experts model with 975 billion total parameters and 41 billion active during processing, trained on 45 trillion tokens across multiple modalities. The model supports a one-million-token context window, enabling extensive reasoning across combined data types, as detailed in the original analysis. Its architecture employs 256 experts, using a combination of global and sliding-window attention, along with specialized modules for images and audio.

Hugging Face reports that running the full model requires approximately 2 terabytes of VRAM for BF16 precision or about 600 GB for NVFP4, making it impractical for most consumer hardware. The model is accessible via supported inference engines, including Transformers, SGLang, vLLM, and llama.cpp, and can also be accessed through hosted inference services. However, details about licensing, training data transparency, and benchmark results are not yet available, highlighting the ongoing need for transparency in large-scale AI models.

At a glance
announcementWhen: announced July 2026
The developmentThinking Machines has released Inkling on Hugging Face, offering a massive multimodal model with significant hardware requirements and limited independent evaluation.

Implications of Inkling’s Open Multimodal Release

The launch of Inkling introduces a highly scalable, open-access multimodal model capable of reasoning across text, images, and audio, which could accelerate research and applications in scientific, media, and enterprise domains. However, its substantial hardware requirements limit direct deployment to organizations with advanced infrastructure, emphasizing reliance on hosted solutions or model quantization.

Amazon

high VRAM graphics card for AI training

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background on Large-Scale Multimodal AI Models

Large language models and multimodal AI systems have rapidly advanced, with models like GPT-4 and PaLM-E demonstrating multimodal capabilities. The release of Inkling by Thinking Machines follows a trend toward open, high-parameter models, although most comparable models remain proprietary or require significant resources. Prior efforts have shown that scaling up parameters enhances performance but also raises accessibility and safety concerns.

“”This model is huge.””

— Hugging Face

Amazon

large-scale AI inference server

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unverified Aspects of Inkling’s Performance and Licensing

There are no independent benchmark results, safety evaluations, or detailed licensing terms available at this stage. The actual performance of Inkling on real-world multimodal tasks, especially video processing, remains untested. It is also unclear whether the model’s open status includes full training data access or usage restrictions.

Amazon

multimodal AI model hardware requirements

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for Testing and Evaluation of Inkling

Developers and research teams are expected to begin testing Inkling through supported inference engines, focusing on latency, accuracy, and resource consumption. Independent evaluations, safety assessments, and domain-specific fine-tuning are anticipated to clarify its capabilities and limitations. Further disclosures on licensing and benchmark results are likely in upcoming updates from Thinking Machines or Hugging Face.

Amazon

AI model deployment hardware

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is Inkling?

Inkling is a large, open multimodal AI model from Thinking Machines that processes text, images, and audio, with 975 billion parameters and a one-million-token context window.

Can I run Inkling on my personal computer?

Likely not. The model requires approximately 2 TB of VRAM for full deployment, making it impractical for typical consumer hardware. Access is expected primarily through hosted inference services or specialized hardware.

Does Inkling support video processing?

While the architecture allows for image inputs with a temporal dimension, native video performance has not been evaluated, and no specific video capabilities are confirmed yet.

What are the licensing terms for Inkling?

The release is described as open, but detailed licensing, usage restrictions, and training data transparency have not been disclosed.

How does Inkling compare to other multimodal models?

Independent benchmark results and safety evaluations are not yet available, so comparisons are premature. Its scale and multimodal capacity are notable, but real-world performance remains to be validated.

Source: ThorstenMeyerAI.com

FLEA & TICK SEAS

Flea & tick season Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The NVIDIA Earnings Preview: What Q1 FY27 Will Reveal About the AI Cycle

NVIDIA reports Q1 FY27 earnings on May 20, 2026, with a revenue forecast of $78 billion, shedding light on the AI cycle and market demand.

ShinyHunters · The New APT Model.

ShinyHunters has evolved into a distributed, AI-enabled extortion collective, marking a shift from traditional APTs. This impacts enterprise security strategies.

Portal By Spotify Cut My Claude Code Token Usage By 90%

Spotify’s Portal platform has reportedly cut Claude Code token consumption by 90%, raising questions about recent platform changes and their impact.

OpenAI Loses Trademark Dispute At EU Court

OpenAI has lost a legal challenge over its trademark at the European Court of Justice, impacting its branding and market strategy in the EU.