Skip to main content
Inkling - Free AI Tool

Inkling

Inkling is a general-purpose multimodal LLM that accepts text, image, and audio inputs to generate text outputs.

No reviews yet
Open Source
What is Inkling?
Inkling is a sophisticated general-purpose multimodal model developed by Thinking Machines and hosted on Hugging Face. It is designed to process diverse input modalities, including text, images, and audio, and subsequently generate text-based outputs. The model is built upon a 66-layer decoder-only transformer architecture, featuring a sparse Mixture-of-Experts (MoE) feed-forward backbone, where each token is routed to 6 of 256 experts, plus 2 shared experts. Its attention mechanism combines local and global layers, contributing to its robust processing capabilities. Inkling is natively multimodal, employing a hierarchical patch encoder for images and video, and discrete token encoding for audio. All these modalities are projected into a shared hidden space, allowing for joint processing by the decoder. With 975 billion total parameters and 41 billion active parameters, it supports BF16 and NVFP4 numerics. Released with open weights, Inkling aims to foster research, fine-tuning, and integration into third-party products. It supports local deployment via popular open-source libraries like SGLang, vLLM, TokenSpeed, Unsloth, and Huggingface, and also offers API access through third-party inference providers. The model is intended for use in English and other languages, as well as across multiple coding languages.
Key Benefits & Features
✓
Multimodal Input Processing

Inkling accepts text input in UTF-8 encoding, image input in any pixel-based format (optimally between 40px and 4096px per dimension), and audio input in WAV format sampled at 16kHz (optimally under 20 minutes in length).

✓
Text Output Generation

The model generates its output exclusively as UTF-8 encoded text, regardless of the input modality.

✓
Open Weights Release

Inkling is released with open weights, enabling researchers and developers to use it for research, fine-tuning, and integration into their own third-party products.

✓
Multilingual and Multi-coding Language Support

The model is intended for use in English and other languages, possessing general multilingual capabilities, and supports multiple coding languages.

✓
Mixture-of-Experts (MoE) Architecture

Inkling utilizes a 66-layer decoder-only transformer architecture with a sparse Mixture-of-Experts (MoE) feed-forward backbone, where each token is routed to 6 of 256 experts, plus 2 shared experts active on every token.

✓
Local Deployment Support

Inkling supports local deployment using several open-source libraries, including SGLang, vLLM, TokenSpeed, Unsloth, and Huggingface's own libraries, with provided recipes for integration.

✓
API Access via Third-Party Providers

Beyond local deployment, Inkling is also accessible via API through various third-party inference providers.

Inkling Pricing
Pricing modelOpen Source
Starting priceFree (open weights)
Free plan—
Free trial—
Billing—

Detailed Pricing Info

Inkling is released with open weights, making it free to use for research, fine-tuning, and integration. While API access is available through third-party inference providers, specific pricing for these services is not provided on the official Inkling page.
Pros & Cons of Inkling
Pros
  • Multimodal capabilities allow processing of text, image, and audio inputs, making it versatile for diverse applications.
  • Released with open weights, fostering research, fine-tuning, and integration into custom applications by downstream developers.
  • Supports a wide range of applications including agentic systems, coding assistants, chatbots, and RAG, suitable for general conversational use and instruction-following.
  • Offers broad language support, including English and general multilingual capabilities, alongside support for multiple coding languages.
  • Supports local deployment with popular open-source libraries like SGLang, vLLM, TokenSpeed, Unsloth, and Huggingface, providing flexibility for developers.
Cons
  • Optimal performance for image inputs requires specific dimensions (40px to 4096px), which might necessitate pre-processing for out-of-range images.
  • Optimal performance for audio inputs is achieved with lengths under 20 minutes, potentially limiting its effectiveness for longer audio processing tasks without segmentation.
  • While strong, Inkling's performance benchmarks indicate it does not consistently outperform all frontier models (both open and closed weights) in every evaluated category, such as Reasoning HLE, SWEBench Pro, and Factuality tasks like AA Omniscience.
  • Safety evaluations identified potential risks, and mitigations were applied before release, suggesting inherent challenges in ensuring safe behavior across multimodal inputs and complex interactions.
Frequently Asked Questions

What is Inkling?

Inkling is a general-purpose multimodal AI model developed by Thinking Machines. It can take text, image, and audio as input and generates text as output. It's designed for various AI applications and is released with open weights.

What types of inputs and outputs does Inkling support?

Inkling accepts text (UTF-8), images (pixel-based, ideally 40px-4096px dimensions), and audio (WAV, 16kHz, ideally under 20 minutes). Its output is always UTF-8 encoded text.

Can Inkling be deployed locally?

Yes, Inkling supports local deployment using several popular open-source libraries, including SGLang, vLLM, TokenSpeed, Unsloth, and Huggingface's own libraries.

What is Inkling's architecture?

Inkling is built on a 66-layer decoder-only transformer architecture. It features a sparse Mixture-of-Experts (MoE) feed-forward backbone, routing each token to 6 of 256 experts plus 2 shared experts, and uses a hybrid of local and global attention layers.
Classification

Related Topics

#Natural Language Processing
#Computer Vision
#Speech Recognition
#Code Generation
#Agentic AI
#Agentic systems
#Tool-use systems
#Coding assistants
#Chatbots
#Retrieval-augmented generation (RAG)
#General-purpose conversational AI
#Instruction-following
#Multimodal tasks
User Reviews & Ratings
(0 reviews)

Write a Review

Community Feedback (0)