Traditional IT monitoring focuses on error rates, availability, and system health. With generative AI applications, however, inappropriate responses and information leaks can occur while the system stays fully operational. You need to continuously check whether interactions with users are appropriate for the business, whether they carry risk, and whether they deliver the expected value.
On top of this, operating generative AI involves many stakeholders, including non-engineers, across business units, operations, risk management, and compliance. Conventional monitoring tools built for engineers cannot support this.
An employee unintentionally enters a customer's personal data or confidential information into a prompt. Citadel Radar detects it at the moment of input and blocks it according to policy.
"Ignore all previous instructions." Citadel Radar detects adversarial input that tries to slip past guardrails and blocks the input before it can trigger an inappropriate response.
Continuously monitors misinformation and harmful or inappropriate output, so you catch changes in performance early.
Comes with built-in metrics designed with reference to AI safety guidelines and various frameworks. You can start monitoring common AI risk areas right away, without building evaluation criteria from scratch.
Use “Block” for risks that must be blocked outright, and “Flag” for risks you want to monitor and analyze. You can choose the control method for each metric based on how the application is used and your organization’s operating policy. Multiple metrics and operating parameters can be managed together as an “evaluation package.”
Stakeholders across business units, risk management, and compliance can check the risk status and trends on a shared screen. The accumulated evaluation metrics and operation logs become a common language for discussing AI risk across the organization.
The architecture was redesigned from the ground up and optimized for firewall workloads. It is a foundation built to sustain fast, automated, always-on monitoring of critical AI systems.
Input evaluation runs in parallel with response generation. Output evaluation achieves low latency through support for local models. A monitoring-only configuration is also available
Maps the relationships between metrics and major frameworks such as Japan's AI Guidelines for Business and OWASP. It lets you evaluate the often-ambiguous notion of AI safety in a structured way.
The violations that Citadel Radar detects in production become assets for preventing recurrence. By connecting them with pre-release evaluation and testing (Citadel Lens), it forms a continuous loop of pre-release testing, production monitoring, and improvement. Even as the generative AI in use changes, it keeps protecting safe AI adoption as a foundation of trust.
LLM evaluation and testing. Quantifies risk with automated and human evaluation.
Pre-deployment review and inventory management of AI assets. AI governance that connects the first and second lines of defense.
Built-in metrics aligned with AI safety guidelines and major industry frameworks provide a structured approach to evaluating AI safety, including criteria based on Japan’s AI Guidelines for Business.
It is a firewall and monitoring product that evaluates the inputs and outputs of generative AI applications and detects, controls, and monitors risk. It continuously evaluates the business appropriateness, risk, and quality of prompts and responses, and brings organization-wide AI governance into day-to-day operation.
It receives input and output logs from your application through an API and evaluates them in real time based on evaluation metrics. It runs evaluation before a response is shown to the user, and depending on the result it can allow the response, block it, or flag it for monitoring.
Use Block for risks you need to stop strictly, such as the input of personal data or serious guideline violations. Use Flag for risks you want to monitor for analysis. You can set this flexibly for each metric based on how the application is used and your organization’s operating policy.
The design keeps the impact to a minimum. Input evaluation runs in parallel with response generation, so it adds almost no latency. Output evaluation also supports lightweight local models, and can achieve low-latency evaluation of around 0.1 seconds.
Yes. Citadel Radar is optimized for monitoring by non-engineers. Stakeholders such as business units, operations, risk management, and compliance can check the risk situation on a shared screen, and can also review decision reasons and violation categories in detail.
On the input side it detects the input of personal or confidential data, prompt injection, jailbreaks, and more. On the output side it detects hallucination, harmful or inappropriate content, and leakage of confidential information. In addition to built-in metrics, you can create custom metrics for your own rules and industry guidelines.
Ready to see Citadel Radar in action?