📊 Full opportunity report: Welcome Inkling By Thinking Machines on ThorstenMeyerAI.com — validation score, market gap, and execution plan.
TL;DR
Thinking Machines has introduced Inkling, a large-scale, open multimodal AI model on Hugging Face. It processes text, images, and audio but requires substantial hardware, and its benchmarks and licensing remain unverified.
Thinking Machines has made available Inkling, a 975-billion-parameter multimodal AI model, on Hugging Face. This development is discussed in What Thinking Machines’ Inkling Signals About AI’s Evolution. This release marks a significant step in open access to large-scale models capable of processing text, images, and audio within a single framework, despite high hardware demands.
Inkling is described as a decoder-only Mixture-of-Experts model with 975 billion total parameters and 41 billion active during processing, trained on 45 trillion tokens across multiple modalities. The model supports a one-million-token context window, enabling extensive reasoning across combined data types, as detailed in the original analysis. Its architecture employs 256 experts, using a combination of global and sliding-window attention, along with specialized modules for images and audio.
Hugging Face reports that running the full model requires approximately 2 terabytes of VRAM for BF16 precision or about 600 GB for NVFP4, making it impractical for most consumer hardware. The model is accessible via supported inference engines, including Transformers, SGLang, vLLM, and llama.cpp, and can also be accessed through hosted inference services. However, details about licensing, training data transparency, and benchmark results are not yet available, highlighting the ongoing need for transparency in large-scale AI models.
Implications of Inkling’s Open Multimodal Release
The launch of Inkling introduces a highly scalable, open-access multimodal model capable of reasoning across text, images, and audio, which could accelerate research and applications in scientific, media, and enterprise domains. However, its substantial hardware requirements limit direct deployment to organizations with advanced infrastructure, emphasizing reliance on hosted solutions or model quantization.

NVIDIA Tesla V100 Volta GPU Accelerator 32GB Graphics Card
Graphics Card Interface: Pci E
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Background on Large-Scale Multimodal AI Models
Large language models and multimodal AI systems have rapidly advanced, with models like GPT-4 and PaLM-E demonstrating multimodal capabilities. The release of Inkling by Thinking Machines follows a trend toward open, high-parameter models, although most comparable models remain proprietary or require significant resources. Prior efforts have shown that scaling up parameters enhances performance but also raises accessibility and safety concerns.
“”This model is huge.””
— Hugging Face

MX3 M.2 AI Accelerator
High-Performance AI Processing: The MX3 is designed to handle the most demanding AI computer vision workloads, delivering exceptional…
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Unverified Aspects of Inkling’s Performance and Licensing
There are no independent benchmark results, safety evaluations, or detailed licensing terms available at this stage. The actual performance of Inkling on real-world multimodal tasks, especially video processing, remains untested. It is also unclear whether the model’s open status includes full training data access or usage restrictions.

Multimodal AI Systems Engineering: Building Production Vision-Language Models, Document AI, and Cross-Modal Retrieval Pipelines (Production AI Engineering Series)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Next Steps for Testing and Evaluation of Inkling
Developers and research teams are expected to begin testing Inkling through supported inference engines, focusing on latency, accuracy, and resource consumption. Independent evaluations, safety assessments, and domain-specific fine-tuning are anticipated to clarify its capabilities and limitations. Further disclosures on licensing and benchmark results are likely in upcoming updates from Thinking Machines or Hugging Face.

Transactions on Large-Scale Data- and Knowledge-Centered Systems XXX: Special Issue on Cloud Computing (Lecture Notes in Computer Science Book 10130)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is Inkling?
Inkling is a large, open multimodal AI model from Thinking Machines that processes text, images, and audio, with 975 billion parameters and a one-million-token context window.
Can I run Inkling on my personal computer?
Likely not. The model requires approximately 2 TB of VRAM for full deployment, making it impractical for typical consumer hardware. Access is expected primarily through hosted inference services or specialized hardware.
Does Inkling support video processing?
While the architecture allows for image inputs with a temporal dimension, native video performance has not been evaluated, and no specific video capabilities are confirmed yet.
What are the licensing terms for Inkling?
The release is described as open, but detailed licensing, usage restrictions, and training data transparency have not been disclosed.
How does Inkling compare to other multimodal models?
Independent benchmark results and safety evaluations are not yet available, so comparisons are premature. Its scale and multimodal capacity are notable, but real-world performance remains to be validated.
Source: ThorstenMeyerAI.com