Interfaze

logo

pricing

help

docs

blog

sign in

The First Open Weight Model for Deterministic Work: Interfaze 1 Lite

copy markdown

Almost every model you can call today was trained to write for a person. interfaze-1-lite is built to be read by your code: open weights, one GPU, and a confidence score on everything it reads.

interfaze-1-lite weights on Hugging Faceinterfaze-1-lite model cardTry interfaze-1-lite in the playground

Models for people, models for machines

A chat model's answer is judged by the person reading it. If it's a little off, the reader notices, rephrases, and moves on.

In a pipeline, the reader is an if statement. It can't tell a confident answer from a lucky guess, so every mistake flows straight into your database.

That changes what a good output looks like:

Built for a personBuilt for a program
OutputProse that reads wellTyped fields that match a schema
LabelsAny word that sounds rightOnly values from the set you defined
Location"Near the top of page 2"A bounding box, a page number, a timestamp
ConfidenceImplied by toneA number on every line, word and field
ConsistencyNice to vary from run to runSame input, same output
MistakesCaught by the readerCaught by a threshold, or not at all

The confidence score is the one that matters most. A model that is right 95% of the time can't automate a task unless it also tells you which answers are in the other 5%.

With a score, the rule is simple: act on anything above your threshold and send the rest to a person. Without one, a person has to check everything, and nothing is automated.

Introducing interfaze-1-lite

interfaze-1-lite is an open-weight, any-to-text model for deterministic work. It takes text, images, audio and files, including PDFs and Word documents, and returns text or JSON that matches your schema.

It's built for tasks where there's one right answer:

  • OCR with a box and confidence for every line and word
  • Speech-to-text with speaker diarization
  • Classification and structured extraction into your schema
  • Object detection and GUI detection
  • Translation, forecasting and guardrails

The weights are Apache 2.0 and the whole model runs on a single 80 GB GPU. Through the API it's $0.85 per million input tokens and $1.50 per million output tokens.

interfaze-1-lite
Input price$0.85 / MTok
Output price$1.50 / MTok
Context window128k tokens
Max output tokens32k tokens
InputText, images, audio, files
OutputText, JSON
ReasoningAvailable
PDF pagesUp to 50 per request
Open weightsYes, runs on one 80 GB GPU
LicenseApache 2.0

The architecture

Lite isn't one network. It's a reasoning core plus specialist models, each built for one kind of perception.

The core reads the request, decides which specialists to run, and writes the answer from what they return. The specialists do the reading, listening and locating that a general model is weakest at.

interfaze-1-lite architecture: inputs flow into a reasoning core that runs specialist models and writes the answer, while each specialist's raw result is returned as precontext

Why split it up

A single vision-language model can read a page well. But it reads the way a person does: it gives you the words, not where they are or how sure it was.

Specialists produce that metadata as part of how they work. A text detector scores every line it finds, and a speech recognizer timestamps every word.

ApproachReads the contentExact positionsConfidence per lineHandles nuance
General VLM aloneYesApproximateNoYes
Specialist models aloneYesYesYesNo
Lite: core + specialistsYesYesYesYes

The core keeps what makes an LLM useful: it understands the request, follows your schema and reasons about what it read. The specialists keep what makes classic models useful: precise output with a score attached.

How the parts combine

OCR stitches two views of one page. The document reader supplies the text in reading order, and the line detector supplies the geometry and confidence.

Each detected line takes the reader's words for it. The box is exact, the text is complete, and the confidence comes with it.

OCR stitching on three rows of a bank statement: the document reader supplies the text, line geometry supplies the boxes and confidence scores 0.86, 0.99 and 0.92, and the stitched lines carry all three

These are real lines and scores from the bank statement in the example below.

The other parts combine the same way:

  • Speakers are attributed per word, by the largest overlap with each speaker's turns, then grouped into chunks.
  • Detection and GUI grounding run on the reasoning core, which returns boxes on a 0 to 1000 grid. Outlines come from the segmentation model.
  • A run task skips planning. Naming one capability runs that specialist directly and returns its raw result.

Everything the specialists produce comes back in precontext, next to the answer. Your code gets the clean JSON and the evidence behind it in the same response.

Benchmarks

Lite is on the leaderboard next to Gemini-3.7-Flash, Claude-Sonnet-5, GPT-5.4-Mini and Grok-4.3. Bold marks where Lite beats all four.

BenchmarkTaskinterfaze-1-liteinterfaze-1Best of the other four
MMMU-ProMultimodal reasoning73.2%71.1%Grok-4.3, 68.7%
RefCOCOVisual grounding83.8%82.1%Gemini-3.7-Flash, 80.9%
SOB value accuracyStructured output81.5%80.5%Gemini-3.7-Flash, 80.2%
olmOCRDocument OCR83.8%85.7%Claude-Sonnet-5, 83.5%
VoxPopuli WER (lower is better)Speech recognition3.0%2.4%Gemini-3.7-Flash, 4.0%
Spider 2.0-LiteText-to-SQL48.9%52.9%Claude-Sonnet-5, 50.6%
GPQA DiamondPhD-level science85.9%92.4%Gemini-3.7-Flash, 91.4%
OCRBench V2Text in images60.9%70.7%Gemini-3.7-Flash, 63.9%
MMMLUMultilingual knowledge87.8%90.9%Grok-4.3, 89.7%

Lite leads on multimodal reasoning, visual grounding, structured output, document OCR and speech. It even edges out our flagship on MMMU-Pro, RefCOCO and SOB.

It trails on text in images, general knowledge and science questions. For those, interfaze-1 scores higher.

View the full benchmark breakdown →

Self-hosting

The weights are on Hugging Face under the Apache 2.0 license. The whole model runs on a single 80 GB GPU with no external services, so it works offline.

You need a GPU with compute capability 8.9 or newer, like an H100, and ffmpeg for audio.

hf download interfaze-ai/interfaze-1-lite requirements.txt --local-dir .
pip install -r requirements.txt

Load it with Transformers and call chat, the same entry point the API uses. The core plans, runs the specialists it needs and answers.

from transformers import AutoModel

model = AutoModel.from_pretrained("interfaze-ai/interfaze-1-lite", trust_remote_code=True)

answer = model.chat(
    [{"role": "user", "content": "What is the total, and which item is highlighted?"}],
    files=["receipt.jpg"],
)
print(answer["content"])

To skip planning, call one capability directly. ocr runs the document reader and line geometry and returns the stitched page.

doc = model.ocr("invoice.pdf", page_range=[1, 2])
print(doc["text"])

The model card covers every capability's method, arguments and output.

Transformers is slower than a batching serving engine. For throughput, use the API.

Use it with the Interfaze API

To try it without code, open the playground and pick interfaze-1-lite from the model picker. It shows the answer and the precontext behind it.

In code, set model to interfaze-1-lite. Everything else works the same as interfaze-1, including structured output and precontext.

Interfaze SDK

Vercel AI SDK

LangChain SDK

import { Interfaze } from "interfaze";

const interfaze = new Interfaze({ apiKey: process.env.INTERFAZE_API_KEY });

const response = await interfaze.chat.completions.create({
	model: "interfaze-1-lite",
	messages: [
		{
			role: "user",
			content: [
				{ type: "text", text: "Extract the details from this ID" },
				{ type: "image_url", image_url: { url: "https://r2public.jigsawstack.com/interfaze/examples/id.jpg" } },
			],
		},
	],
});

console.log(response.choices[0]?.message.content);

Interfaze speaks the Chat Completions standard, so any OpenAI-compatible SDK works too. Point it at https://api.interfaze.ai/v1 with your key from the dashboard.

Example: read a bank statement, classify every row

Here's the kind of job Lite is built for. A brokerage statement comes in as an image, and your bookkeeping system needs every account activity row as a typed record with a category.

A quarterly brokerage statement from Northbridge Investments with an account summary, holdings, and an account activity table

The plan:

  1. Extract the account details and every activity row into a schema.
  2. Classify each row into a fixed set of categories, defined as an enum so no other label can come back.
  3. Gate each row on the OCR confidence of the line it came from. Rows above 0.9 post automatically, the rest go to a person.

Interfaze SDK

Vercel AI SDK

LangChain SDK

import { Interfaze, responseFormat } from "interfaze";
import { z } from "zod";

const interfaze = new Interfaze({ apiKey: process.env.INTERFAZE_API_KEY });

const StatementSchema = z.object({
	account_number: z.string(),
	account_holder: z.string(),
	statement_period: z.string(),
	transactions: z.array(
		z.object({
			date: z.string(),
			description: z.string(),
			amount: z.number(),
			category: z.enum(["contribution", "withdrawal", "dividend", "interest", "capital_gain", "balance"]),
		})
	),
});

const response = await interfaze.chat.completions.create({
	model: "interfaze-1-lite",
	messages: [
		{
			role: "user",
			content: [
				{ type: "text", text: "Extract the account details and every Account Activity row, and classify each row." },
				{
					type: "image_url",
					image_url: { url: "https://r2public.jigsawstack.com/interfaze/examples/financial_document_example.png" },
				},
			],
		},
	],
	response_format: responseFormat(z.toJSONSchema(StatementSchema), "bank_statement"),
});

type OcrLine = { text: string; average_confidence: number };
type OcrResult = { sections: { lines: OcrLine[] }[] };

const statement = StatementSchema.parse(JSON.parse(response.choices[0]?.message.content ?? "{}"));
const ocr = response.precontext?.find((p) => p.name === "ocr")?.result as OcrResult | undefined;
const lines = ocr?.sections.flatMap((section) => section.lines) ?? [];

const REVIEW_BELOW = 0.9;

for (const row of statement.transactions) {
	const line = lines.find((l) => l.text.startsWith(row.date));
	const confidence = line?.average_confidence ?? 0;
	const status = confidence >= REVIEW_BELOW ? "auto" : "review";
	console.log(status, row.date, row.category, row.amount, confidence);
}

The output

The schema response has every row, typed and classified. The withdrawal even keeps its minus sign.

{
  "account_number": "NB-2468-1357",
  "account_holder": "MICHAEL T. ANDERSON",
  "statement_period": "April 1, 2024 - June 30, 2024",
  "transactions": [
    { "date": "04/01/24", "description": "Beginning Market Value", "amount": 1125780.31, "category": "balance" },
    { "date": "04/15/24", "description": "Contribution", "amount": 10000, "category": "contribution" },
    { "date": "04/22/24", "description": "DIVIDEND RECEIVED - VANGUARD S&P 500 ETF", "amount": 185.67, "category": "dividend" },
    { "date": "05/03/24", "description": "WITHDRAWAL - ACH -", "amount": -5000, "category": "withdrawal" },
    { "date": "05/15/24", "description": "Contribution", "amount": 5000, "category": "contribution" },
    { "date": "05/24/24", "description": "INTEREST RECEIVED - CORE POSITION", "amount": 23.14, "category": "interest" },
    { "date": "06/03/24", "description": "DIVIDEND RECEIVED - ISHARES MSCI EAFE ETF", "amount": 132.98, "category": "dividend" },
    { "date": "06/17/24", "description": "Contribution", "amount": 10000, "category": "contribution" },
    { "date": "06/28/24", "description": "CAPITAL GAIN DISTRIBUTION - VTI", "amount": 286.35, "category": "capital_gain" },
    { "date": "06/30/24", "description": "Ending Market Value", "amount": 1209509.65, "category": "balance" }
  ]
}

Precontext carries the OCR behind it, with a box and confidence for every line and word. Here it is trimmed to the withdrawal line, without the per-word boxes:

{
  "name": "ocr",
  "result": {
    "extracted_text": "Account Summary\nCustomer Service\n...",
    "sections": [
      {
        "lines": [
          {
            "text": "05/03/24 WITHDRAWAL - ACH -",
            "bounds": {
              "top_left": { "x": 31, "y": 774 },
              "top_right": { "x": 198, "y": 774 },
              "bottom_right": { "x": 198, "y": 786 },
              "bottom_left": { "x": 31, "y": 786 }
            },
            "average_confidence": 0.86,
            "words": [
              { "text": "05/03/24", "confidence": 0.99 },
              { "text": "WITHDRAWAL", "confidence": 0.99 },
              { "text": "-", "confidence": 0.92 },
              { "text": "ACH", "confidence": 0.99 },
              { "text": "-", "confidence": 0.42 }
            ]
          }
        ]
      }
    ]
  }
}

And the gate prints:

auto 04/01/24 balance 1125780.31 0.99
auto 04/15/24 contribution 10000 0.98
auto 04/22/24 dividend 185.67 0.95
review 05/03/24 withdrawal -5000 0.86
auto 05/15/24 contribution 5000 0.99
auto 05/24/24 interest 23.14 0.92
auto 06/03/24 dividend 132.98 0.95
auto 06/17/24 contribution 10000 0.98
auto 06/28/24 capital_gain 286.35 0.98
auto 06/30/24 balance 1209509.65 0.99

Nine rows post on their own. The withdrawal goes to a person, because the trailing dash after "ACH" read at 0.42 and pulled the line down to 0.86.

The extraction was right, and a reviewer confirms it in seconds. That's the point: the model told you exactly which row to look at, and the box tells the reviewer exactly where on the page to look.

A chat model would have returned the same JSON with no way to tell which row was shaky. Here, the category can't fall outside your enum, and every row carries the confidence of the text it came from.

Interfaze

logo

Product

Playground

OCR

Models

Leaderboards

Pricing

OpenWebSearch

DefaultModel