Phocinae Largha 150M V1
Phocinae Largha 150M V1 by Phocinae, a text-classification model. Understand and compare features, benchmarks, and capabilities.
Comparison
| Feature | Phocinae Largha 150M V1 | Interfaze |
|---|---|---|
| Input Modalities | text | image, text, audio, video, document |
| Native OCR | No | Yes |
| Long Document Processing | No | Yes |
| Language Support | unknown | 162+ |
| Native Speech-to-Text | No | Yes |
| Native Object Detection | No | Yes |
| Guardrail Controls | Yes | Yes |
| Context Input Size | unknown | 1M |
| Tool Calling | Yes | Tool calling supported + built in browser, code execution and web search |
Scaling
| Feature | Phocinae Largha 150M V1 | Interfaze |
|---|---|---|
| Scaling | Self-hosted/Provider-hosted with quantization | Unlimited |
View model card on Hugging Face
斑海豹 Largha, the spotted seal: a 144.3M decision model (en-first; zh evaluated on machine-translated cases) (150M-class) for structured decisions — one forward pass per decision, on your own hardware. Not a chat model: it takes a state plus a list of typed questions (noul yes/no · choice pick-one · score 2–10) and returns calibrated answers with confidence, robust to option reordering. Downloads: Hugging Face · GitHub · ModelScope 魔搭 · 中文: FAQ 中文版 · 场景演示画廊.
TL;DR
| TL;DR | value |
|---|---|
| parameters | 144.3M (mmBERT-small base: hidden 384 × 22 layers × 6 heads, 256k-vocab tokenizer, RoPE + sliding-window + full attention; base context 8192; decision sequence ≤512 tokens with a 192-token head attention window) |
| typed-decisions en (400 cases / 2000 decisions) | 0.906 (specialist: fitted on this dataset's train split) — Laya 0.766 (self-measured, native interface) · JEV 0.727 (generalist, zero-shot) · meraGPT 0.768 (generalist, zero-shot) |
| typed-decisions zh (machine-translated cases; training mix includes machine-translated Chinese ≈2,400 rows and native Chinese ≈1,400 rows) | 0.848 |
| option-order flip robustness (lower is better) | CPU fp32: flip150 0.0200 · flip400 0.0217 · random-mean 0.0144 · any 0.0283 (GPU fp16 0.0200/0.0217, note only) |
| inference latency | GPU fp16 p50 21.0 ms (RTX 5090) · CPU single-thread p50 1.64 s per case (1 case = 1 state + 5 questions, single forward pass) · CPU 8-thread batch 8–20 decisions/s |
| JevBench public-231 | 0.5455 (126/231) — below the 58.4% gate, disclosed honestly; tool_selection 12/12 (n=12) |
| escalate routing (E1 gate, τ=0.6) | +0.0876 kept-subset acc (0.906→0.9936 (kept subset)) while −55.0% LLM calls (79.6% at τ=0.5) (45.0% official-set escalate; τ swept on the eval set — re-scan per domain) |
| calibration | shipped column ECE 0.2519 (en) / 0.1941 (zh); with the bundled calibration column: 0.0168 / 0.0152; calibration temperatures 0.8660205/0.8081192/0.6624661 (applied at inference) |
Full numbers, methodology, and evidence: BENCHMARKS.md.
See it in action
📽️ 全部 27 个场景演示 → docs/gallery/(审批安全 · 路由省费 · 实时分级 · 办公文档 · 流程工程 · 对照与可靠性;中文版 gallery/README_cn.md)
Why Phocinae: slash agent costs
55.0% fewer LLM calls, 21.0 ms per decision on an RTX 5090. Phocinae-Largha-150M (144.3M params) routes the repetitive decisions inside agent sessions — command approvals, tool selection (12/12 on JevBench tool_selection, k≤10), step checks, output screening — to a local single-forward-pass engine (8192-token base context, 512-token default head) instead of a 500–4,000-token API call. The τ=0.6 confidence gate cuts LLM traffic 55.0% with kept-subset accuracy 0.906 → 0.9936 and full routed accuracy 0.8135; deterministic inference means auditable, replayable decisions with zero data leaving your machine — at 21.0 ms GPU p50 (RTX 5090) vs 1.5 s+ API round-trips. Full cost model: docs/cost-savings.md.
Quick start — phocinae-server
phocinae-serverships the runtime, not the weights — fetch the weights first (step 1).
The official runtime is phocinae-server: a local FastAPI service (127.0.0.1 only) that loads these weights with a pure PyTorch forward (no extra runtime needed). Full spec: docs/protocol.md. Hardware tiers: docs/deployment.md.
hf download Phocinae/Phocinae-Largha-150M-v1 --local-dir ./largha
pip install phocinae-server
PHOC_MODEL_DIR=./largha python -m phocinae.main # http://127.0.0.1:8155The directory produced by that download is exactly the layout PHOC_MODEL_DIR
expects (model.safetensors, encoder/, tokenizer/, rl_agent_config.json).
From Python (the same engine the server uses):
from phocinae.engine import Engine
eng = Engine("./largha", device="cpu") # device="auto" selects CUDA when present
answers, confidence, action, usage = eng.run(
"The agent restarted nginx after checking the logs and the health endpoint is green.",
[{"id": "ok", "type": "noul"},
{"id": "act", "type": "choice", "options": ["allow", "ask", "deny"]},
{"id": "risk", "type": "score"}],
)One decision over HTTP:
curl -s http://127.0.0.1:8155/v1/systemone -H 'Content-Type: application/json' -d '{
"model": "Phocinae-Largha-150M-v1",
"state": "The agent restarted nginx after checking the logs and the health endpoint is green.",
"questions": [
{"id": "ok", "type": "noul", "threshold": 0.65},
{"id": "act", "type": "choice", "options": ["allow", "ask", "deny"]},
{"id": "risk", "type": "score"}
]
}'{
"model": "Phocinae-Largha-150M-v1",
"answers": {"ok": false, "act": 0, "risk": 3},
"usage": {"input_tokens": 146, "output_tokens": 0},
"answer_confidence": {"ok": 0.668, "act": 0.409, "risk": 0.1633},
"action": {"act": {"act_probability": 0.1581}},
"routing": {"model": "Phocinae-Largha-150M-v1", "device": "cpu", "perm": "none", "backend": "phocinae-pure-torch"}
}Known limitation in tool routing.
Router.route_tool()confuses semantically close tool names. Reproducible case: state "The user wants to find the latest news about the product launch." with the six toolsweb_search / read_file / run_command / list_files / fetch_url / ask_userreturnsfetch_url, notweb_search, at confidence ≈0.33. Treat low-confidence tool picks as escalate-worthy rather than final.
answers values: noul = bool · choice = 0-based option index · score = integer 2–10. Errors: 422 (unknown model/type, >64 questions, choice without options, score with options, threshold outside [0,1]) · 413 (>2 MiB body) · 401 (bearer token). Extension keys answer_confidence / action / routing can be disabled with PHOC_EXTENSIONS=0.
Documentation
- BENCHMARKS.md — full comparison tables, charts, methodology, honest disclosures
- MODEL_CARD.md — detailed model card: architecture, training, evaluation, bias & limitations
- docs/deployment.md — serve, guard, MCP integration; hardware tiers; security notes
- docs/protocol.md — the
/v1/systemonedecision protocol - docs/reproduce.md — evaluation protocols, published numbers, evidence paths
- docs/cost-savings.md — LLM-cost model for the escalate gate
- docs/gallery/ — 27 application scenarios with animated demos (中文: gallery/README_cn.md)
- calib/ — optional post-hoc calibration column (power transform; ECE 0.2519/0.1941 → 0.0168/0.0152)
- docs/faq.md — common questions (中文: faq.zh.md) · docs/technical-report.md — short technical report
What it is / what it is not
- For: structured decisions — approval gates, tool routing, escalation, document triage, step checks, output screening; anywhere you want a fast, cheap, local decision layer instead of a large model.
- Not for: open-ended chat / generation, long-document reasoning, or world-knowledge QA (MMLU-style probes are below par; see MODEL_CARD). Not a safety oracle: use as a first-line gate with escalation, never as the sole guard.
Honest disclosures
- JevBench public-231: 0.5455 (126/231) vs a 58.4% acceptance gate — not passed. We publish the number as measured, and we never train on the eval rows.
- zh results are on machine-translated cases; the training mix includes machine-translated Chinese (≈2,400 rows) and native Chinese (≈1,400 rows) — treat zh as an in-mix (fitted) evaluation, not zero-shot cross-lingual transfer.
- Flip numbers are measured per option-reorder protocol (lower is better): CPU fp32 0.0200/0.0217 (flip150/400 reversed) · random-mean 0.0144 · any 0.0283. GPU fp16 0.0200/0.0217 and 1k-row 0.0181/0.0150/0.0331 are note-only values from different protocols.
- Calibration: the shipped column has ECE 0.2519 (en). Calibration temperatures (0.8660205/0.8081192/0.6624661) are stored in the model repo config and applied at inference by phocinae-server. A recommended recalibration column is bundled under
calib/.
Weights & license
model.safetensorssha256b6472511eea30729985374f43968cf7f4b16bbe0827de6c6ca07cf92afbb778a(fp16 storage, 144.3M params); trained fromjhu-clsp/mmBERT-smallonLocalLLaMA/typed-decisions(train split + flip-augmented reorderings), recipe inrl_agent_config.json.- Apache-2.0 (see LICENSE); base encoder
jhu-clsp/mmBERT-smallis MIT (see NOTICE). · Repo layout:model.safetensors·encoder/·tokenizer/·rl_agent_config.json·checkpoint_meta.json·figures/·docs/.
Credits
- typed-decisions (Apache-2.0, LocalLLaMA HF org) — protocol & test data
- mmBERT-small (JHU CLSP) — base encoder
- JevBench — held-out protocol used for disclosure
Version & machine-readable sources
- This release:
v1.1— pinned tag on the model repository (v1.0.0marks the previous release), for reproducible citation. Resolve its commit withgit ls-remote --tags https://huggingface.co/Phocinae/Phocinae-Largha-150M-v1. (No commit hash is written into this file on purpose: it would go stale on the next commit.) - Machine-readable facts (same numbers as this card, for AI systems and retrieval pipelines): llms.txt
- Citation metadata: CITATION.cff
- Landing page (mirrors this card, with structured data): https://phocinae.github.io/Phocinae-Largha-150M-v1/
Citation
Machine-readable: CITATION.cff. BibTeX:
@misc{phocinae2026largha,
title = {Phocinae-Largha-150M-v1: A 144M-parameter bilingual typed decision model},
author = {Phocinae Project},
year = {2026},
url = {https://huggingface.co/Phocinae/Phocinae-Largha-150M-v1},
license = {Apache-2.0}
}Revision history
- 2026-10-09 — v1.1 refresh. Weights upgraded (each metric in this card re-measured on the new weights; previous release sha256
db79d5ee2f16597f34e564f5a4363bddb5b5bbd9827c01819725dabcc7802697). Figures/gallery re-rendered; optional calibration column added undercalib/. - 2026-10-09 — docs. Expanded quick start (weight download, Python & HTTP examples, tool-routing caveat); added landing page & machine-readable sources section.