
Work Here?
Fastino offers Task-Specific Language Models (TLMs) for enterprises. These models perform targeted tasks like document summarization, JSON conversion, PII redaction, and function calling, running on CPUs/NPUs for near-instant inference without GPUs. It differentiates through task-focused optimization, CPU-based inference, and flat-rate pricing, with deployment in VPC, on-premise, or edge for data security. Its goal is to provide efficient, reliable AI tools that fit enterprise workflows with strong data security and cost efficiency.
Industries
Data & Analytics
Enterprise Software
AI & Machine Learning
Company Size
11-50
Company Stage
Seed
Total Funding
$24.5M
Headquarters
Palo Alto, California
Founded
2024
See people who can refer or advise you
Help us improve and share your feedback! Did you find this helpful?
Total Funding
$24.5M
Above
Industry Average
Funded Over
2 Rounds
Industry standards
Fastino releases GLiNER2.5: A boundary-prediction architecture that removes span enumeration from Information extraction. August 24, 2026 Information extraction teams face a recurring choice. Small encoder models are cheap but rigid, and large language models are flexible but expensive per document. Fastino released GLiNER2.5 to narrow that gap. The release replaces span enumeration with boundary prediction: the model scores where an entity starts and ends instead of scoring every candidate span against a width grid. That single change removes the maximum entity width, allows a 4,096-word context, and keeps computation linear in sequence length for a fixed schema. It also unlocks joint entity-relation decoding, cross-task label constraints, and per-span attributes. Across 16 zero-shot benchmarks, the multilingual checkpoint reaches 56.17 overall macro F1 against 56.09 for GLiNER2, with a 24.75-point gain on XNLI. Three checkpoints ship on Hugging Face under Apache 2.0 at 74M, 194M, and 287M parameters. Is it deployable? Yes, Fastino released three GLiNER2.5 checkpoints on Hugging Face under Apache 2.0, with local inference on CPU, CUDA, or MPS via pip install "gliner2[local]" (Python 3.10+). No inference provider currently hosts the checkpoints, so self-hosting is the deployment path. * Company level: any tier. The 74M and 194M checkpoints run on standard CPU boxes, so a two-person team can ship extraction without GPU budget. Larger orgs get a fine-tunable, privately hosted alternative to per-token LLM extraction. * Industries: legal and contract operations, healthcare and clinical documentation, financial services, insurance claims, customer support, and AI safety tooling. * Applications: PII detection and redaction, contract clause extraction, knowledge graphs for agent memory, agent and model routing, guardrail classification, clinical entity extraction with negation and dosage attributes. What changed. Earlier GLiNER models located entities by enumerating candidate spans: every start position paired with every allowed width, each scored against the schema. That design tied compute to a width axis and imposed a hard ceiling on entity length. GLiNER2.5 removes enumeration. The shared encoder still processes text and schema queries in one pass. Instead of scoring spans, the model predicts start and end scores over token boundaries plus inside scores over tokens. A sparse proposal stage selects the most promising starts and ends per query and pairs them, with no restriction on distance. A reranking head then scores each candidate using boundary evidence and span content. Relation candidates are drawn from the same pool rather than a separate path. Fastino team reports that computation stays linear in sequence length for a fixed schema and candidate budget. Five capabilities that follow. * Long-context extraction: Removing explicit span representations cut memory enough to train on sequences up to 4,096 words. The checkpoints ship with max_len=4096. The library also adds native chunking helpers (extract_entities_long, extract_long, Classifier.classify_long, JointIE.extract_long) that remap spans to character offsets in the original document. * Unlimited span length: GLiNER2 enumerated spans up to a fixed width, typically around twelve words; longer entities were never scored. In GLiNER2.5 a span can open at the first token and close at the last. A forty-word indemnification clause costs the same to locate as a two-word name. * Joint entity and relation extraction: Users declare entity types, typed relations, and structural rules (unique_head=True, no_self_loops, and a beam search assembles a globally consistent graph. Invalid combinations are never admitted, so output conforms by construction. Check result.feasible before using the graph. * Constrained classification: C.implies and C.excludes rules bind labels across tasks during decoding. Fastino's own GLiGuard guardrail model illustrates the problem being solved: without constraints, a prompt can be labeled safe while simultaneously flagged for prompt injection. If no valid assignment exists, the classifier raises an error. * Span attributes: Attribute groups such as sentiment attach to specific entity types via applies_to, and are decoded span-by-span in the same forward pass. Entities return qualified rather than flat. The model family. All three share the same public API. Load with AutoExtractor, not the legacy GLiNER2 span loader. Benchmarks. Fastino team evaluates zero-shot on 16 public datasets, reporting macro F1 against GLiNER2 at matched sizes. Overall average: GLiNER2.5 Multi reaches 56.17 versus 56.09 for GLiNER2 Multi. GLiNER2.5 Base reaches 54.87 versus 53.34. The headline gain is XNLI, where Multi jumps to 62.30 from 37.55, a 24.75-point increase. Few-NERD improves for Base to 55.14 from 47.22. Romanian RONEC, an untrained language, improves for both. Key takeaways. * Boundary prediction replaces span enumeration; entity width no longer costs compute. * Three Apache 2.0 checkpoints: 74M, 194M, 287M, all CPU-runnable. * Joint decoding returns schema-valid graphs, removing post-hoc validation layers. * Overall F1 rises to 56.17 (Multi) and 54.87 (Base); extraction average dips for Multi. * Chunking keeps a span only when both boundaries land in one chunk.
Fastino Labs releases open-weight finance and healthcare models. August 12, 2026 Fastino Labs released two open-weight models for finance and healthcare, post-trained entirely by an autonomous agent on NVIDIA's Nemotron 3.5 Lightning base. On 11 August 2026, Palo Alto-based applied AI research lab Fastino Labs (creators of the GLiNER model family) announced two domain-specific open-weight models: Fastino-Nemotron-3.5-Lightning-Finance and Fastino-Nemotron-3.5-Lightning-Healthcare. Both models were post-trained on NVIDIA's newly released Nemotron 3.5 Lightning, a 30-billion parameter Mixture-of-Experts model with 3 billion active parameters. Both models are available immediately on Hugging Face under the permissive Apache 2.0 open-source licence. Both models were post-trained entirely using the Fastino Fine-Tuning Agent, an autonomous autoresearch agent that handles task research, data curation, evaluation set creation, parallel training runs, error recovery, data contamination tests, and final checkpoint selection from plain language prompts. The agent completed the full post-training workflow for both models in less than a day, a process that typically requires weeks for human post-training teams. The Fastino Fine-Tuning Agent is available as a private preview and will see a full general release supporting various open-weight models in the coming weeks. For the finance model, FinQA execution accuracy lifted performance from 15.86% (base model) to 59.23% (+43.37 percentage points), while TAT-QA (F1) increased from 19.01% to 56.63% (+37.62 percentage points). It outperformed the base model across five distinct financial reasoning benchmarks (including SEC-Num, FinEntity, and BizFinBench). Fastino Labs' healthcare model demonstrated gains across eight medical benchmarks, including verified improvements on MEDEC, MedCalc-Bench, and HealthBench (improving HealthBench from 40.62% to 48.17% across 700 reserved clinical conversations). Both models demonstrated capability transfer to unseen related tasks, confirming generalised domain proficiency rather than benchmark overfitting.
Fastino Labs has released two open-weight AI models specialised for finance and healthcare, post-trained on NVIDIA's Nemotron 3.5 Lightning. The models were developed entirely by the company's Fine-Tuning Agent, an autonomous system that completed the work in under 10 hours per model. The finance model improved FinQA accuracy from 15.86% to 59.23%, whilst the healthcare model achieved verified gains across eight medical benchmarks. Both 30-billion-parameter models are available on Hugging Face under Apache 2.0 licence. The Fine-Tuning Agent, now in private preview, autonomously handles the entire post-training pipeline including data curation, evaluation set creation, and testing for data contamination. Fastino Labs, founded in 2024, has raised $25 million from investors including Khosla Ventures and Insight Partners.
GLiNER2-PII: open source privacy filtering with PII detection. By: Mary Newhauser & Urchade Zaratiana May 14, 2026 Fastino, Inc. is releasing GLiNER2-PII, a 300M parameter SOTA open-source model for detecting and redacting PII. Today Fastino, Inc. is releasing GLiNER2-PII, a 300M parameter state-of-the-art open-source model that outperforms OpenAI's Privacy Filter, NVIDIA-PII and other leading PII models on accuracy. GLiNER2-PII is a 300 million parameter multilingual model for detecting and redacting personally identifiable information in unstructured text. It is designed for production privacy workflows, supports a fine-grained taxonomy of 42 entity types out of the box, and can be adapted at inference time to any custom schema without retraining. On the SPY benchmark, GLiNER2-PII achieves the highest span-level F1 of any publicly available PII model, outperforming four leading models including OpenAI's Privacy Filter. Every text document that flows through a modern software system is a potential privacy incident. Names, addresses, account numbers, and credentials are embedded across support tickets, healthcare records, and financial documents. This unstructured data requires detection and redaction before it can safely move through a pipeline. Whether the downstream consumer is an enterprise compliance pipeline or an agentic AI system like Hermes Agent or OpenClaw acting on a user's behalf, sensitive information must be identified before it goes any further. Building systems that do this reliably is genuinely difficult. PII spans are heterogeneous, locale-dependent, and frequently ambiguous without context. Existing approaches fall short in different ways: decoder-based models like OpenAI's Privacy Filter repurpose a 1.5B parameter autoregressive checkpoint as a token classifier, but lock developers into a fixed schema of 8 entity types that cannot be customized at inference time. Other open-source models fall short on accuracy. GLiNER2-PII takes a different approach. The model is label-conditioned, meaning the target schema is an input to the model, not a property baked into its weights. This lets the same checkpoint serve any organization's PII policy without retraining, whether that means broad masking for analytics pipelines or fine-grained redaction for compliance audits. The core breakthrough was in the post-training data. Real PII annotations are inherently unshareable, which has historically constrained every open PII model to small, narrow, or synthetic datasets of questionable quality. In post-training of GLiNER2-PII, Fastino, Inc. used Pioneer's synthetic data agent to generate 4,910 high-quality annotated examples across seven languages and a wide range of document formats: chat logs, support tickets, CRM notes, KYC forms, invoices, and medical records. The result is the most accurate PII model available, released under a permissive license for both research and production use. Detect, extract, and redact 42 types of sensitive information in a single pass. * Personal identity: person, full name, first name, middle name, last name, date of birth * Contact and location: email, phone number, address, street address, city, state or region, postal code, country * Government and tax identifiers: government ID, national ID number, passport number, driver's license number, license number, tax ID, tax number * Banking and payment: bank account, account number, routing number, IBAN, payment card, card number, card expiry, card CVV * Digital identity: username, IP address, account ID, sensitive account ID * Secrets and credentials: password, secret, API key, access token, recovery code * Sensitive dates: sensitive date, document date, expiration date, transaction date Because GLiNER2-PII is an encoder model, all 42 entity types are evaluated simultaneously in a single forward pass, producing deterministic outputs with no sampling variance or hallucinated spans. OpenAI's Privacy Filter, by comparison, repurposes a 1.5B parameter decoder checkpoint as a token classifier locked to 8 fixed entity types. GLiNER2-PII offers more than 5x the label coverage at a fraction of the memory footprint, with a schema that can be customized at inference time without retraining. This granularity matters in practice. Rather than flagging a broad "financial information" category, the model distinguishes between a card number, its expiry, and its CVV as separate entities, and can identify an API key embedded inside a URL or a recovery code in a support ticket. This is what allows downstream systems to apply differentiated redaction policies, such as retaining a card expiry for transaction records while fully masking the CVV. How Fastino, Inc. evaluated GLiNER2-PII. Fastino, Inc. evaluated on the SPY (Synthetic PII Yesterday) benchmark, which contains 200 documents split evenly between legal Q&A forums and medical transcripts, annotated with seven PII types. Fastino, Inc. chose SPY because it provides recent, naturally formatted text that none of the models in its comparison were trained on, making it a genuine out-of-distribution test. Fastino, Inc. compared GLiNER2-PII against four publicly available PII detectors: OpenAI's Privacy Filter, GLiNER PII (NVIDIA), gliner_multi_pii-v1 (urchade), and gliner-pii-base-v1.0 (Knowledgator). These represent a range of approaches, from OpenAI's repurposed decoder to several GLiNER-based extractors, and each uses a different internal label set, so Fastino, Inc. applied an identical deterministic label mapping across all five systems to ensure a fair comparison. Best-in-class PII detection. On the SPY benchmark, GLiNER2-PII achieved the highest overall accuracy of any system Fastino, Inc. tested: * Average F1 of 0.471, vs. 0.391 (NVIDIA GLiNER PII), 0.384 (urchade), 0.373 (OpenAI Privacy Filter), and 0.368 (knowledgator). * Recall of 0.722 on legal documents and 0.681 on medical documents, the highest of any model tested. * OpenAI's Privacy Filter achieved comparable recall (0.640 and 0.671) but with precision of just 0.250 and 0.271, compared to GLiNER2-PII's 0.354 and 0.355. * Performance was consistent across both the legal and medical domains. Recall is the metric that matters most in redaction workflows, because a missed entity means sensitive data goes unmasked. By this measure, GLiNER2-PII and OpenAI's Privacy Filter are the only two systems that catch the majority of PII spans, but OpenAI's model gets there by flagging far more false positives. GLiNER2-PII finds more sensitive information while making fewer incorrect predictions, which in practice means less noise for downstream systems to deal with. The consistency across domains is also worth noting: the model was trained entirely on synthetic data generated by Pioneer, yet it generalizes effectively to naturally occurring legal and medical text it has never seen. How Fastino, Inc. trained GLiNER2-PII. Training a high-quality PII detection model requires two things: a robust base architecture and a large, diverse corpus of annotated examples. For the architecture, Fastino, Inc. fine-tuned GLiNER2, its 0.3B parameter multi-task encoder model for structured information extraction. For the data, Fastino, Inc. turned to Pioneer. Data generation. The central challenge in training PII models is that the data you need is exactly the data you cannot freely collect, share, or annotate. Real PII is sensitive by definition, and the datasets that do exist publicly are either too small, too narrow in scope, or already used as training data by other models, making honest evaluation difficult. Fastino, Inc. addressed this using Pioneer's constraint-driven synthetic data generation pipeline, which takes a label schema and a natural language description of the target task and produces diverse, fully annotated training examples without ever touching real personal data. The pipeline takes two inputs: the list of 42 entity types and a natural language description of the task. From these, it builds a set of rules that control what each generated example looks like. Some rules are structural, ensuring that each example contains a specific mix of entity types and that coverage across the full label set is balanced. Others control variety, specifying things like document format, language, and writing style. For each example, a subset of these rules is sampled and used to prompt multiple large language models, which produce the text along with annotations. This process generated 4,910 examples spanning seven languages and a wide range of document formats, from chat logs and support tickets to invoices and medical records. Conclusion. As AI agents gain access to personal data and the ability to act on it, the risk surface for PII leakage has expanded significantly, and detection needs to keep pace. Production deployments require models that are both flexible enough to adapt to different compliance requirements and specific enough to distinguish between closely related entity types. GLiNER2-PII meets both of these requirements: 42 fine-grained entity types, zero-shot customization, and deterministic inference in a single forward pass. On the SPY benchmark, it achieved the highest F1 of any system Fastino, Inc. evaluated, outperforming OpenAI's Privacy Filter, NVIDIA's GLiNER PII, and two other publicly available detectors. Its recall scores, 0.722 on legal documents and 0.681 on medical documents, were also the highest in the comparison, which matters because in redaction, every missed entity is sensitive data left exposed. GLiNER2-PII is now available as an open source model under the Apache 2.0 license, continuing the tradition of its GLiNER models and contributing to the broader effort to make reliable privacy tooling accessible to everyone. Try GLiNER2-PII in Pioneer. GLiNER2-PII is now available on for inference on Pioneer.
Fastino, a startup based in Palo Alto, has raised $17.5 million. The company specializes in training small AI models designed for specific tasks, which it sells to businesses.
Find jobs on Simplify and start your career today
Industries
Data & Analytics
Enterprise Software
AI & Machine Learning
Company Size
11-50
Company Stage
Seed
Total Funding
$24.5M
Headquarters
Palo Alto, California
Founded
2024
Find jobs on Simplify and start your career today