
Work Here?
Langfuse provides open-source observability and analytics for Language Learning Machine (LLM) applications. It helps teams trace and debug complex LLM workflows by recording production traces that show every step in an LLM chain, including unlimited nested actions, exact costs, and token-level details for prompts and completions. It also tracks non-LLM actions like database queries and API calls to give a full view of an application's performance. The platform offers prebuilt analytics focused on token usage, cost, latency, and trace scores, all linked to traces to help identify root causes quickly. Langfuse integrates with popular frameworks and libraries and offers a public API and typed SDKs for Python and JS/TS so teams can build custom dashboards and features. Its goal is to make it easier to observe, cost, and optimize LLM-powered applications by providing clear visibility and actionable insights.
Industries
Data & Analytics
Enterprise Software
AI & Machine Learning
Company Size
11-50
Company Stage
Seed
Total Funding
$4.1M
Headquarters
Berlin, Germany
Founded
2022
See people who can refer or advise you
Help us improve and share your feedback! Did you find this helpful?
Total Funding
$4.1M
Below
Industry Average
Funded Over
2 Rounds
Industry standards
Remote Work Options
Hybrid Work Options
Unlimited Paid Time Off
Paid Vacation
Flexible Work Hours
Wellness Program
Mental Health Support
Conference Attendance Budget
Professional Development Budget
Stock Options
Company Equity
401(k) Retirement Plan
401(k) Company Match
Healthcare?
Phone/Internet Stipend
Home Office Stipend
Jev-as-a-judge in Langfuse Evaluators. Jev-based evals are now available in Langfuse, so you can score production traffic at up to 40 to 400x lower cost than LLM-based approaches. Jev-based evals are now available in Langfuse. Jev-like models give you and your team a way to run evals on production traffic at scale and identify signals from agent executions and user interactions. All of this at up to 40 to 400x lower cost than with LLM-based approaches. Jev is part of a new model category that prioritizes performance on one dimension over general ability of the model. It can be seen as a "zero-shot classifier". As founder Diogo Almeida said on the latest Latent Space episode, they optimize for "Intelligence per Dollar" and design the model with fast, structured, machine-native decision-making inside software in mind. While the category is being formed, Langfuse is excited to bring Jev-based evals into Langfuse already today. Jev, TypeSafe's new model, handles narrow, typed questions over unstructured data. It is best suited for quick decisions that are repeated, high volume, and the possible answers are known before the call. Jev offers three different output types: * Choice picks one option from a set you define, up to 255, and returns the probability of each plus a confidence value. * Score rates the state against ordered rubric levels and returns a probability-weighted value, the full distribution, and a confidence value. Up to 10 levels. * Noul answers yes or no and returns the probability it is true. It carries no separate confidence field, so code that reads answer.confidence on everything will break on binaries. Jev's key strengths align well with best practices for evaluating your AI application, while overcoming some key issues of LLM-as-a-judge setups. Jev allows you to run evals on production traces at scale. The input tokens are up to 400x cheaper than frontier models. Where sampling was previously the strategy for cost efficient production monitoring, the full traffic can now be evaluated. The same state can be used across multiple questions. All questions are handled concurrently and independently. Where previously context creep and interaction effects of multiple LLM-based judgments interfered, Jev now treats it as separate assessments. The predefined output structures force you to think in distinct decisions during setup. You define a clear atomic question and in the output options need to specify the conditions for each verdict. As TypeSafe says, Jev is judging what you say, not what you mean. While execution speed is not a bottleneck in async production evals, this new model category also allows for significantly faster execution. Setting up Jev-based evals is possible via the Evaluators tab in the Langfuse app.
Arcjet launches agent runtime security for production AI agents. New product helps teams discover agent activity, apply controls before and after consequential actions, and preserve evidence for security reviews and compliance SAN FRANCISCO, Sept. 17, 2026 /PRNewswire/ - Arcjet, the security platform that ships in your AI code, today launched agent runtime security, a new product that helps engineering teams secure the AI agents they are building while giving security teams the governance and compliance evidence they need. Arcjet brings observability, enforcement, and audit capabilities across agent workflows so teams can discover which agents are running, control what they can do, and understand what happened and why. AI agents are moving beyond chat interfaces and into production workflows, where they can read and write to databases, respond to support tickets, refund payments, call tools and APIs, and take other actions on behalf of users. Those workflows can start from a chat interface, an email, a text message, a code commit, or another event, and can continue autonomously across multiple systems. As agents take on longer-running workflows, security teams need to answer three questions across the full sequence of activity, which agents are running, whether a particular action should be allowed, and what happened and why. Arcjet's agent runtime security addresses those questions through observe, enforce, and audit capabilities. Teams can discover agent activity and connect actions across sessions, apply deterministic security policies before and after calls to LLMs, tools, databases, and APIs, and preserve the execution context needed for security reviews and compliance. "Agents are now taking real actions inside production systems, which means security teams need to know which agents are operating and what they have done, and apply controls at machine speed," said David Mytton, CEO at Arcjet. "A risky outcome can develop across a series of steps that look perfectly reasonable on their own. Arcjet connects those steps and gives teams policy controls to detect them." Arcjet's agent runtime security centers on three parts of securing agents in production, observe, enforce, and audit. Observe: Discover all your agents Arcjet supports ingesting agent activity without application code changes or deploying another agent. Platform and security teams can use existing OpenTelemetry observability tooling to send activity directly to Arcjet for real-time visualization and analysis. For teams using Claude, Arcjet can also pull activity from the Claude Compliance API. Arcjet connects activity across sessions so teams can see an agent's sequence of actions as one workflow rather than a collection of unrelated events. Activity can include prompts, tool call parameters, session metadata, identity, security decisions, and other application context, giving teams a view of what each agent is doing across a run. Agent identity and inventory are part of that visibility. Arcjet gives teams an inventory of the agents and applications operating inside their environment, with activity and individual runs associated with each agent so teams can inspect actions and security decisions step by step. Enforce: Apply controls before and after every action Once teams can see their agents and activity across sessions, Arcjet lets security teams define controls for prompt injection detection, PII and sensitive information leak prevention and redaction, automation and bot detection, rate limits, and quota controls. Arcjet guards apply deterministic policies to tools, APIs, database calls, and other inputs and outputs. Powered by Rego and Open Policy Agent, teams can create versioned, immutable policies through Arcjet's web UI, API, CLI, or MCP without redeploying application code. Policies can define the actions an agent is allowed to take, such as restricting recipients or attachments in an email tool, setting acceptable bounds for refund values, or limiting web fetch tools to trusted API URLs. Arcjet returns a decision to the application before the action executes, allowing the application to stop the operation, request human approval, or return an explanation to the agent. Applied before and after calls to LLMs, tools, databases, and APIs, these controls can mitigate risk before consequential actions and verify results before the workflow continues. For enforcement, Arcjet has native integrations with major agent frameworks, including Claude Agents SDK, Claude Managed Agents, OpenAI Agents SDK, LangChain, LangFuse, Strands, Mastra, and Microsoft's Agent Framework. This in-code context allows Arcjet to track recorded actions, their inputs, and policy decisions across the workflow. Audit: Evidence and proof of compliance Arcjet collects the context of each execution so teams can reconstruct what happened, understand why a policy decision was made, and provide evidence for security reviews and compliance audits. Correlated traces preserve actions, inputs, security decisions, and policy evaluations across the work SOURCE Arcjet
Our goal continues to be building the best LLM engineering platform
ClickHouse valued at $15 billion, acquires Langfuse in AI analytics market. ClickHouse, a database analytics firm, has raised $400 million in a series D round led by Dragoneer, valuing the company at $15 billion. The investment highlights the growing demand for real-time analytics software as companies deploy AI features. ClickHouse competes with Databricks and Snowflake, and has also acquired Langfuse, an open-source platform for creating and testing large language models. Ask Aime: How might the recent investment in ClickHouse impact its stock performance and potential for growth in the real-time analytics software market? Or continue with others.
ClickHouse, a database management company, has raised $400 million in a Series D round led by Dragoneer Investment Group, valuing the firm at $15 billion. Bessemer Venture Partners, GIC and Index Ventures also participated. The funding reflects investor enthusiasm for real-time analytics providers as companies deploying AI tools require fast data management for soaring volumes. ClickHouse's technology enables rapid query responses, competing with Databricks and Snowflake in helping firms build and run AI applications. Founded in 2009, ClickHouse serves customers including Meta, Cursor, Sony and Tesla with open-source database software for real-time analytics. The company also announced the acquisition of Langfuse, an open-source platform for creating and monitoring large language models. The valuation follows Databricks' recent $4 billion raise at $134 billion valuation.
Find jobs on Simplify and start your career today
Industries
Data & Analytics
Enterprise Software
AI & Machine Learning
Company Size
11-50
Company Stage
Seed
Total Funding
$4.1M
Headquarters
Berlin, Germany
Founded
2022
Find jobs on Simplify and start your career today