Interfaze

logo

pricing

help

docs

blog

sign in

All models

Phocinae Largha 150M V1

Phocinae Largha 150M V1 by Phocinae, a text-classification model. Understand and compare features, benchmarks, and capabilities.

Comparison

FeaturePhocinae Largha 150M V1Interfaze
Input Modalities

text

image, text, audio, video, document

Native OCRNoYes
Long Document ProcessingNoYes
Language Support

unknown

162+

Native Speech-to-TextNoYes
Native Object DetectionNoYes
Guardrail ControlsYesYes
Context Input Size

unknown

1M

Tool CallingYes

Tool calling supported + built in browser, code execution and web search

Scaling

FeaturePhocinae Largha 150M V1Interfaze
Scaling

Self-hosted/Provider-hosted with quantization

Unlimited

View model card on Hugging Face

斑海豹 Largha, the spotted seal: a 144.3M decision model (en-first; zh evaluated on machine-translated cases) (150M-class) for structured decisions — one forward pass per decision, on your own hardware. Not a chat model: it takes a state plus a list of typed questions (noul yes/no · choice pick-one · score 2–10) and returns calibrated answers with confidence, robust to option reordering. Downloads: Hugging Face · GitHub · ModelScope 魔搭 · 中文: FAQ 中文版 · 场景演示画廊.

TL;DR

TL;DRvalue
parameters144.3M (mmBERT-small base: hidden 384 × 22 layers × 6 heads, 256k-vocab tokenizer, RoPE + sliding-window + full attention; base context 8192; decision sequence ≤512 tokens with a 192-token head attention window)
typed-decisions en (400 cases / 2000 decisions)0.906 (specialist: fitted on this dataset's train split) — Laya 0.766 (self-measured, native interface) · JEV 0.727 (generalist, zero-shot) · meraGPT 0.768 (generalist, zero-shot)
typed-decisions zh (machine-translated cases; training mix includes machine-translated Chinese ≈2,400 rows and native Chinese ≈1,400 rows)0.848
option-order flip robustness (lower is better)CPU fp32: flip150 0.0200 · flip400 0.0217 · random-mean 0.0144 · any 0.0283 (GPU fp16 0.0200/0.0217, note only)
inference latencyGPU fp16 p50 21.0 ms (RTX 5090) · CPU single-thread p50 1.64 s per case (1 case = 1 state + 5 questions, single forward pass) · CPU 8-thread batch 8–20 decisions/s
JevBench public-2310.5455 (126/231) — below the 58.4% gate, disclosed honestly; tool_selection 12/12 (n=12)
escalate routing (E1 gate, τ=0.6)+0.0876 kept-subset acc (0.906→0.9936 (kept subset)) while −55.0% LLM calls (79.6% at τ=0.5) (45.0% official-set escalate; τ swept on the eval set — re-scan per domain)
calibrationshipped column ECE 0.2519 (en) / 0.1941 (zh); with the bundled calibration column: 0.0168 / 0.0152; calibration temperatures 0.8660205/0.8081192/0.6624661 (applied at inference)

Full numbers, methodology, and evidence: BENCHMARKS.md.

See it in action

📽️ 全部 27 个场景演示 → docs/gallery/(审批安全 · 路由省费 · 实时分级 · 办公文档 · 流程工程 · 对照与可靠性;中文版 gallery/README_cn.md)

Why Phocinae: slash agent costs

55.0% fewer LLM calls, 21.0 ms per decision on an RTX 5090. Phocinae-Largha-150M (144.3M params) routes the repetitive decisions inside agent sessions — command approvals, tool selection (12/12 on JevBench tool_selection, k≤10), step checks, output screening — to a local single-forward-pass engine (8192-token base context, 512-token default head) instead of a 500–4,000-token API call. The τ=0.6 confidence gate cuts LLM traffic 55.0% with kept-subset accuracy 0.906 → 0.9936 and full routed accuracy 0.8135; deterministic inference means auditable, replayable decisions with zero data leaving your machine — at 21.0 ms GPU p50 (RTX 5090) vs 1.5 s+ API round-trips. Full cost model: docs/cost-savings.md.

Quick start — phocinae-server

phocinae-server ships the runtime, not the weights — fetch the weights first (step 1).

The official runtime is phocinae-server: a local FastAPI service (127.0.0.1 only) that loads these weights with a pure PyTorch forward (no extra runtime needed). Full spec: docs/protocol.md. Hardware tiers: docs/deployment.md.

hf download Phocinae/Phocinae-Largha-150M-v1 --local-dir ./largha




pip install phocinae-server
PHOC_MODEL_DIR=./largha python -m phocinae.main   # http://127.0.0.1:8155

The directory produced by that download is exactly the layout PHOC_MODEL_DIR expects (model.safetensors, encoder/, tokenizer/, rl_agent_config.json).

From Python (the same engine the server uses):

from phocinae.engine import Engine

eng = Engine("./largha", device="cpu")   # device="auto" selects CUDA when present
answers, confidence, action, usage = eng.run(
    "The agent restarted nginx after checking the logs and the health endpoint is green.",
    [{"id": "ok",   "type": "noul"},
     {"id": "act",  "type": "choice", "options": ["allow", "ask", "deny"]},
     {"id": "risk", "type": "score"}],
)

One decision over HTTP:

curl -s http://127.0.0.1:8155/v1/systemone -H 'Content-Type: application/json' -d '{
  "model": "Phocinae-Largha-150M-v1",
  "state": "The agent restarted nginx after checking the logs and the health endpoint is green.",
  "questions": [
    {"id": "ok",   "type": "noul",   "threshold": 0.65},
    {"id": "act",  "type": "choice", "options": ["allow", "ask", "deny"]},
    {"id": "risk", "type": "score"}
  ]
}'
{
  "model": "Phocinae-Largha-150M-v1",
  "answers": {"ok": false, "act": 0, "risk": 3},
  "usage": {"input_tokens": 146, "output_tokens": 0},
  "answer_confidence": {"ok": 0.668, "act": 0.409, "risk": 0.1633},
  "action": {"act": {"act_probability": 0.1581}},
  "routing": {"model": "Phocinae-Largha-150M-v1", "device": "cpu", "perm": "none", "backend": "phocinae-pure-torch"}
}

Known limitation in tool routing. Router.route_tool() confuses semantically close tool names. Reproducible case: state "The user wants to find the latest news about the product launch." with the six tools web_search / read_file / run_command / list_files / fetch_url / ask_user returns fetch_url, not web_search, at confidence ≈0.33. Treat low-confidence tool picks as escalate-worthy rather than final.

answers values: noul = bool · choice = 0-based option index · score = integer 2–10. Errors: 422 (unknown model/type, >64 questions, choice without options, score with options, threshold outside [0,1]) · 413 (>2 MiB body) · 401 (bearer token). Extension keys answer_confidence / action / routing can be disabled with PHOC_EXTENSIONS=0.

Documentation

  • BENCHMARKS.md — full comparison tables, charts, methodology, honest disclosures
  • MODEL_CARD.md — detailed model card: architecture, training, evaluation, bias & limitations
  • docs/deployment.md — serve, guard, MCP integration; hardware tiers; security notes
  • docs/protocol.md — the /v1/systemone decision protocol
  • docs/reproduce.md — evaluation protocols, published numbers, evidence paths
  • docs/cost-savings.md — LLM-cost model for the escalate gate
  • docs/gallery/ — 27 application scenarios with animated demos (中文: gallery/README_cn.md)
  • calib/ — optional post-hoc calibration column (power transform; ECE 0.2519/0.1941 → 0.0168/0.0152)
  • docs/faq.md — common questions (中文: faq.zh.md) · docs/technical-report.md — short technical report

What it is / what it is not

  • For: structured decisions — approval gates, tool routing, escalation, document triage, step checks, output screening; anywhere you want a fast, cheap, local decision layer instead of a large model.
  • Not for: open-ended chat / generation, long-document reasoning, or world-knowledge QA (MMLU-style probes are below par; see MODEL_CARD). Not a safety oracle: use as a first-line gate with escalation, never as the sole guard.

Honest disclosures

  • JevBench public-231: 0.5455 (126/231) vs a 58.4% acceptance gate — not passed. We publish the number as measured, and we never train on the eval rows.
  • zh results are on machine-translated cases; the training mix includes machine-translated Chinese (≈2,400 rows) and native Chinese (≈1,400 rows) — treat zh as an in-mix (fitted) evaluation, not zero-shot cross-lingual transfer.
  • Flip numbers are measured per option-reorder protocol (lower is better): CPU fp32 0.0200/0.0217 (flip150/400 reversed) · random-mean 0.0144 · any 0.0283. GPU fp16 0.0200/0.0217 and 1k-row 0.0181/0.0150/0.0331 are note-only values from different protocols.
  • Calibration: the shipped column has ECE 0.2519 (en). Calibration temperatures (0.8660205/0.8081192/0.6624661) are stored in the model repo config and applied at inference by phocinae-server. A recommended recalibration column is bundled under calib/.

Weights & license

  • model.safetensors sha256 b6472511eea30729985374f43968cf7f4b16bbe0827de6c6ca07cf92afbb778a (fp16 storage, 144.3M params); trained from jhu-clsp/mmBERT-small on LocalLLaMA/typed-decisions (train split + flip-augmented reorderings), recipe in rl_agent_config.json.
  • Apache-2.0 (see LICENSE); base encoder jhu-clsp/mmBERT-small is MIT (see NOTICE). · Repo layout: model.safetensors · encoder/ · tokenizer/ · rl_agent_config.json · checkpoint_meta.json · figures/ · docs/.

Credits

Version & machine-readable sources

  • This release: v1.1 — pinned tag on the model repository (v1.0.0 marks the previous release), for reproducible citation. Resolve its commit with git ls-remote --tags https://huggingface.co/Phocinae/Phocinae-Largha-150M-v1. (No commit hash is written into this file on purpose: it would go stale on the next commit.)
  • Machine-readable facts (same numbers as this card, for AI systems and retrieval pipelines): llms.txt
  • Citation metadata: CITATION.cff
  • Landing page (mirrors this card, with structured data): https://phocinae.github.io/Phocinae-Largha-150M-v1/

Citation

Machine-readable: CITATION.cff. BibTeX:

@misc{phocinae2026largha,
  title  = {Phocinae-Largha-150M-v1: A 144M-parameter bilingual typed decision model},
  author = {Phocinae Project},
  year   = {2026},
  url    = {https://huggingface.co/Phocinae/Phocinae-Largha-150M-v1},
  license = {Apache-2.0}
}

Revision history

  • 2026-10-09 — v1.1 refresh. Weights upgraded (each metric in this card re-measured on the new weights; previous release sha256 db79d5ee2f16597f34e564f5a4363bddb5b5bbd9827c01819725dabcc7802697). Figures/gallery re-rendered; optional calibration column added under calib/.
  • 2026-10-09 — docs. Expanded quick start (weight download, Python & HTTP examples, tool-routing caveat); added landing page & machine-readable sources section.

Want more deterministic results?

Interfaze

logo

Product

Playground

OCR

Models

Leaderboards

Pricing

OpenWebSearch

DefaultModel