Citadel Radar

Protect your GenAI applications
with a real-time AI firewall.

Challenges

Traditional IT system monitoring is not enough
to protect generative AI applications

Traditional IT monitoring focuses on error rates, availability, and system health. With generative AI applications, however, inappropriate responses and information leaks can occur while the system stays fully operational. You need to continuously check whether interactions with users are appropriate for the business, whether they carry risk, and whether they deliver the expected value.

On top of this, operating generative AI involves many stakeholders, including non-engineers, across business units, operations, risk management, and compliance. Conventional monitoring tools built for engineers cannot support this.

Risk Scenarios

Catch the incidents that happen
while your system looks perfectly healthy

32 privacy

Input of confidential
or personal data

An employee unintentionally enters a customer's personal data or confidential information into a prompt. Citadel Radar detects it at the moment of input and blocks it according to policy.

32 warning sign

Prompt injection and
jailbreak attempts

"Ignore all previous instructions." Citadel Radar detects adversarial input that tries to slip past guardrails and blocks the input before it can trigger an inappropriate response.

32 line chart

Hallucination and
harmful outputs

Continuously monitors misinformation and harmful or inappropriate output, so you catch changes in performance early.

Capabilities

Define your risks, control them flexibly,
and monitor in real time

Define risk with metrics tied to AI safety guidelines
Evaluation categories & metrics mapped to major frameworks

01 — Define

Define risk with metrics
tied to AI safety guidelines

Comes with built-in metrics designed with reference to AI safety guidelines and various frameworks. You can start monitoring common AI risk areas right away, without building evaluation criteria from scratch.

  • Real-time evaluation of input and output logs through an API
  • Support for Japanese-language business context
  • Custom metrics for your own rules, industry guidelines, and compliance requirements
  • A playground to check metric accuracy in advance
MONITOR | Organization-wide monitoring that non-engineers can use
Evaluation metric settings — Input/Output × Block/Monitoring

02 — Control

Flexible control by use case
with Block and Flag

Use “Block” for risks that must be blocked outright, and “Flag” for risks you want to monitor and analyze. You can choose the control method for each metric based on how the application is used and your organization’s operating policy. Multiple metrics and operating parameters can be managed together as an “evaluation package.”

  • Run evaluation before a response is shown, blocking risky output before it reaches the user
  • Review the violation category and reason for blocked inputs and outputs in detail
Monitor Organization-wide monitoring that non-engineers can use
Visually intuitive dashboard

03 — Monitor

Organization-wide monitoring that non-engineers can use

Stakeholders across business units, risk management, and compliance can check the risk status and trends on a shared screen. The accumulated evaluation metrics and operation logs become a common language for discussing AI risk across the organization.

  • Check individual logs, decision reasons, and metadata on the traffic page
  • Visualize detection trends, Block and Flag counts, and risk distribution in analytics
  • Threshold-based email alerts (immediate or scheduled)
  • Per-application API key management, with operational control through role- and application-based access controls

Architecture

A design built for always-on, automated monitoring

High fault tolerance
and scalability

The architecture was redesigned from the ground up and optimized for firewall workloads. It is a foundation built to sustain fast, automated, always-on monitoring of critical AI systems.

Minimal impact on
response time

Input evaluation runs in parallel with response generation. Output evaluation achieves low latency through support for local models. A monitoring-only configuration is also available

An evaluation framework
tied to guidelines

Maps the relationships between metrics and major frameworks such as Japan's AI Guidelines for Business and OWASP. It lets you evaluate the often-ambiguous notion of AI safety in a structured way.

AI Trust Cycle

Feed production insights back into test design

The violations that Citadel Radar detects in production become assets for preventing recurrence. By connecting them with pre-release evaluation and testing (Citadel Lens), it forms a continuous loop of pre-release testing, production monitoring, and improvement. Even as the generative AI in use changes, it keeps protecting safe AI adoption as a foundation of trust.

Citadel Lens Evaluation

LLM evaluation and testing. Quantifies risk with automated and human evaluation.

Citadel Lens Governance

Pre-deployment review and inventory management of AI assets. AI governance that connects the first and second lines of defense.

AI Trust Cycle. From pre-deployment review to production monitoring. Govern the entire lifecycle

Standards

Risk detection aligned with
major guidelines and frameworks

Built-in metrics aligned with AI safety guidelines and major industry frameworks provide a structured approach to evaluating AI safety, including criteria based on Japan’s AI Guidelines for Business.

FAQs

It is a firewall and monitoring product that evaluates the inputs and outputs of generative AI applications and detects, controls, and monitors risk. It continuously evaluates the business appropriateness, risk, and quality of prompts and responses, and brings organization-wide AI governance into day-to-day operation.

It receives input and output logs from your application through an API and evaluates them in real time based on evaluation metrics. It runs evaluation before a response is shown to the user, and depending on the result it can allow the response, block it, or flag it for monitoring.

Use Block for risks you need to stop strictly, such as the input of personal data or serious guideline violations. Use Flag for risks you want to monitor for analysis. You can set this flexibly for each metric based on how the application is used and your organization’s operating policy.

The design keeps the impact to a minimum. Input evaluation runs in parallel with response generation, so it adds almost no latency. Output evaluation also supports lightweight local models, and can achieve low-latency evaluation of around 0.1 seconds.

Yes. Citadel Radar is optimized for monitoring by non-engineers. Stakeholders such as business units, operations, risk management, and compliance can check the risk situation on a shared screen, and can also review decision reasons and violation categories in detail.

On the input side it detects the input of personal or confidential data, prompt injection, jailbreaks, and more. On the output side it detects hallucination, harmful or inappropriate content, and leakage of confidential information. In addition to built-in metrics, you can create custom metrics for your own rules and industry guidelines.

Add real-time guardrails on generative AI in production

Ready to see Citadel Radar in action?