Evaluate and govern your AI systems to minimize AI risk and maximize AI performance.
Generative AI is a probabilistic system that can produce different outputs from the same input. Inappropriate outputs caused by hallucination, harmful responses, and jailbreaks cannot be caught by conventional testing methods, and evaluating them takes deep expertise and considerable effort. On top of this, generative AI and RAG applications update their reference data frequently, so verifying quality by hand every time is not realistic.
Compliance with international regulations and standards, such as Japan’s AI Guidelines for Business and the EU AI Act, is also essential. Organizations need a way to quantify risk against objective criteria and maintain supporting evidence and a clear audit trail.
Citadel Lens comes with built-in automated metrics for risks specific to generative AI, including source-to-answer consistency, question-to-answer relevance, toxicity, and jailbreak detection. The Lens Safety Dataset spans 11 categories, such as safety, fairness, intellectual property, and personal data leakage, and delivers comprehensive, high-detection evaluation suited to multilingual contexts.
Public-facing chatbots are exposed to attacks such as jailbreaks every day. Citadel Lens includes automated red teaming that tests your application from an attacker’s perspective, so you can carry out security validation without relying on a specialist team.
Generative AI risks must be assessed in the context of each application’s intended use. Custom Metrics let you build your own evaluation criteria easily from templates, and human annotation reflects frontline knowledge. Together they enable a double check that combines automated evaluation with human judgment (human-in-the-loop).
In addition to generative AI applications, Citadel Lens supports quality evaluation of machine learning models such as image recognition. It automatically runs tests such as prediction visualization and explanation, robustness stress testing, and detection of bias in models and datasets, and it generates ISO-compliant reports. It is used by BSI (British Standards Institution) for EU AI Act conformity assessments and for the safety evaluation of AI.
Automatically generates synthetic test data with distortions from camera, environment, and position (such as noise and contrast changes), and measures how robust a model is against issues like image degradation.
Automatically detects data segments with high error rates and surfaces bias in models and datasets, so you can see at a glance the segments in which model accuracy drops.
Along with technical validation reports that speed up quality improvement, Citadel Lens also provides compliance assessment reports against AI-related laws, regulations, and international standards. It can serve as a shared platform for the engineers who build AI applications and for compliance teams.
Citadel Lens consists mainly of two components: Evaluation and Governance. Evaluation handles pre-deployment assessment and testing, while Governance standardizes review workflows and manages the AI inventory. Citadel Radar detects and controls risks in production. Detection results and risk trends from Citadel Radar’s production monitoring feed back into Evaluation’s test design, creating a continuous loop of pre-deployment testing, production monitoring, and improvement that supports the entire generative AI lifecycle end to end.
Centralizes AI deployment review, risk assessment, and inventory management, and operationalizes enterprise-wide AI governance.
AI firewall and monitoring for production. Detects and controls risk in real time.
Tokio Marine Holdings: The first full-scale adoption of AI governance with Citadel Lens in the financial industry. Across its domestic group, it makes risk visible from AI model planning through operation and monitors quality automatically.
ONO PHARMA: The first full-scale deployment of Citadel Lens in Japan’s pharmaceutical industry. Lens continuously validates the quality of generative AI used across the company, helping maintain response accuracy, identify risks early, and create a safer environment for AI adoption.
BSI: Following a year-long technical evaluation, BSI selected Citadel AI from among 54 companies worldwide and adopted its technology as a technical and regulatory assessment tool for high-risk AI systems.
Citadel Lens is a platform that quantitatively evaluates and tests the quality and risk of generative AI (LLM) models and machine learning models such as image recognition. Its proprietary technology combines automated and human evaluation to make risks like hallucination, harmful responses, and jailbreaks visible before release, and to keep monitoring them after release. With Citadel Lens Governance, you can also seamlessly run pre-deployment review processes and manage AI asset inventories through an intuitive and simple UI.
It comes with multilingual built-in metrics for source-to-answer consistency (hallucination), question-to-answer relevance, toxicity, jailbreak detection, and more. The Lens Safety Dataset, made up of 11 categories such as safety, fairness, intellectual property, and personal data leakage, enables comprehensive risk evaluation.
Yes. It supports evaluation in Japanese and other languages. It includes a domestically built safety dataset that accounts for Japanese nuance and Japan-specific business practices and cultural context, enabling risk management tailored to Japanese-language use cases.
Yes. The Custom Metrics feature lets you build your own evaluation criteria easily from templates. Even abstract criteria that are hard to express in code can be defined with a prompt and run automatically as quantitative metrics.
Along with automated evaluation, Citadel Lens provides human annotation (human-in-the-loop). It reflects frontline knowledge through a small amount of annotation and double-checks the automated results, a design that raises the reliability of evaluation.
Yes. It includes automated red teaming features. Through jailbreak prompts, generation of adversarial datasets, and data augmentation such as synonym and attribute substitution, you can run security and robustness validation yourself.
No. Citadel Lens applies unified, general-purpose testing that does not depend on the AI model or application. Along with SaaS, it supports deployment on customer-managed cloud environments (IaaS) and on-premises environments.
Ready to see Citadel Lens in action?