Building an In-House LLM Red Teaming Team: Organizing and Operating Security Testing

Introduction
Building an in-house structure for LLM red teaming means establishing a system in which the security, AI, and business functions work together—through clear role division and a combination of automation and human judgment—to continuously detect and remediate vulnerabilities in LLMs that are already in production. This article is intended for security leaders, AI operations teams, and information security departments at companies that have already deployed LLMs, and explains how to design such a structure based on standards like the NIST AI RMF and the OWASP LLM Top 10. By the end, you will understand how to design a role-division chart and testing process suited to your organization, along with concrete steps for running a quarterly implementation cycle. Note that this article presents general guidance for structure building; for individual implementation decisions, we recommend consulting with the supervising expert or specialist.
AI red teaming does not end with a single assessment conducted before deployment. Because vulnerabilities change in response to model updates and prompt changes made during production operation, missed detections will accumulate without a continuous structure in place.
Why an Internal Structure Is Needed: Operational Challenges After Tool Adoption (Comparison Table)
Even after introducing automated diagnostic tools such as Promptfoo, if an operational structure is not established, detection results tend to be left unaddressed, and vulnerabilities end up going unmanaged after all. The scope of issues that can be handled differs clearly between operating a tool on its own and operating it as part of an in-house structure.
| Challenge | Situation with Tool-Only Deployment | Situation After Building an In-House Structure |
|---|---|---|
| Scope of vulnerability detection | Only automated scanning of known patterns | Automated diagnostics plus human judgment also evaluate jailbreaks and context-dependent vulnerabilities |
| Response after detection | Results depend on individual judgment of staff and tend to be left unaddressed | A remediation flow including business-department approval is defined |
| Retesting after prompt changes | No mechanism to detect changes, so retesting is missed | An established process runs retesting triggered by changes |
| Information sharing across departments | Security and AI/ML departments each grasp the situation separately | Responsibilities and reporting lines are clear based on RACI |
| Continuity | Ends with a one-time assessment at the time of introduction | Improvements accumulate through a quarterly cycle |
As the table shows, introducing a tool is merely a means of expanding detection scope; without a response flow after detection and cross-departmental coordination, its effect is limited. In particular, keeping pace with prompt changes and model updates cannot be sustained through operations that rely on the goodwill of individual staff. When judging the priority of building such a structure, use as your first criterion whether your organization is in a state of "detection is happening, but response is not keeping up."
Basic Structure of an LLM Red Team: Required Roles and Departments
An LLM red team structure cannot be completed by a single department; it functions only when three departments—security, AI/ML, and business—each bring a different axis of judgment to the table. Technical vulnerability assessment and business-impact assessment are separate matters, and clarifying from the outset which department is responsible for what is the starting point of structure design.
Role of the Security Department: Test Planning and Result Evaluation
The security department is responsible for planning tests and making the final judgment on evaluating results. Specifically, using the OWASP LLM Top 10 categories (prompt injection, sensitive information disclosure, excessive agency, etc.) as a basis, it decides which categories will be the target of a given round of testing, and translates the scope of target systems, exclusion items, and implementation period into a test plan. This planning stage also includes receiving model configuration information and system prompt content provided by the AI/ML department, and verifying the validity of attack scenarios. When narrowing down test targets, clearly specifying in the plan which items can be verified with a single input—such as direct injection—and which require multi-step verification via external data sources—such as indirect injection—helps prevent the implementation period estimate from going off track.
In evaluating results, detected vulnerabilities must be classified by severity, distinguishing between known vulnerabilities registered in external databases such as CVE and context-dependent issues specific to LLMs (such as jailbreak success rates and the reach of indirect injection). When using a metric such as Attack Success Rate (ASR), a key practical point of evaluation is to continuously measure the same attack patterns in a form that allows comparison before and after model updates.
Furthermore, treating the output of automated diagnostic tools such as Promptfoo and PyRIT as a first-pass screening, and then manually removing false positives and assigning priority, also falls within the security department's scope of responsibility. The severity and priority determined here become the input information for the subsequent approval process carried out by the business department. By handling everything consistently from test planning through to results evaluation, the security department can make clear the basis for the improvement measures that the AI/ML department implements.
Role of the AI/ML Department: Understanding Model Behavior and Implementing Improvements
The AI/ML department is responsible for explaining, from the model's operating principles, why a detected vulnerability occurred, and for implementing improvement measures. For the security department to assess the validity of an attack scenario, technical information such as the configuration of the system prompt, the scope of fine-tuning training, and RAG's chunk size and retrieval settings is required, and the AI/ML department serves as the starting point for providing this information. For example, in a case where a jailbreak succeeded, the department reproduces how the constraints of the system prompt were circumvented through rephrasing of the input, and determines whether the cause lies in the ordering of the instruction text or in the distribution of the training data.
In implementing improvements, multiple countermeasures are considered, such as revising the system prompt, adding guardrail rules, and strengthening RAG's grounding checks, and the department evaluates whether the fix would compromise the quality of existing responses. Since fixes involving model updates or fine-tuning incur higher verification costs than ordinary feature improvements, determining in advance the extent to which the issue can be addressed through prompt-side fixes becomes a practical decision criterion.
After the fix, the department retests using the same attack pattern and shares with the security department whether the attack success rate has dropped to an acceptable range. In systems that use RAG, there are also vulnerabilities originating from the external data source itself, such as RAG poisoning, and in such cases it becomes necessary to go as far as reviewing the input management of the vector database and the chunk generation process. The practice of keeping records of the scope of implemented improvements and verification results can also be used when deciding whether to retest at the time of model updates.
Role of the Business Department: Risk Prioritization and Approval
The business department is responsible for overlaying the severity of a vulnerability as assessed by the security department with a separate axis—business impact—to determine priorities, and for approving the implementation of improvement measures. Even when the technical attack success rate is high, the urgency of response differs depending on whether the affected feature is used infrequently and the resulting business loss is small, or whether the success rate is moderate but occurs in a feature with many users, such as a customer-facing chatbot. This judgment cannot be made by the security department or the AI/ML department alone; it requires the business department's judgment, which has an understanding of contract terms and the actual state of customer interactions.
During the approval process, the department also confirms the impact that implementing the improvement measure would have on existing response quality and response speed. For example, if adding a guardrail rule overly restricts responses to legitimate inquiries, the business department decides whether to prioritize addressing the vulnerability or user satisfaction, and determines whether to approve or hold the release. As a hypothetical example, if a vulnerability involving the leakage of confidential information is found in a RAG-based internal inquiry response system, and the fix lowers the response accuracy for some legitimate inquiries, the department compares the number of affected users against the damage caused by an information leak, and decides whether to approve the release of the fixed version.
When making such decisions, it is also important to document in advance, not as numerical values but as conditional statements, the upper limit of risk that would be acceptable if a response were put on hold. For example, predefining a condition such as "immediately hold if a response containing personal information occurs" ensures that judgment does not waver between departments even in an emergency. Furthermore, it is also the business department's role to review the cumulative state of vulnerabilities detected during each quarterly test cycle and decide on the budget and expanded test scope for the next cycle. This record of approval also serves as the basis for operating the subsequent roles and responsibilities table.
Red Team Role Assignment Table (RACI)
In RACI, it is important to distinguish between Responsible and Accountable. By consolidating final decision-making into a single person or department while distributing hands-on work among specialists in each area, you can avoid situations where work stalls while waiting for approval.
| Work Item | Security Department | AI/ML Department | Business Department |
|---|---|---|---|
| Formulating the test plan | AR | C | I |
| Providing technical information on attack scenarios | C | AR | I |
| Detecting vulnerabilities and assessing severity | AR | C | I |
| Root cause analysis and implementing improvement measures | C | AR | I |
| Assessing business impact and prioritization | I | I | AR |
| Approving the release of improvement measures | C | C | AR |
| Retesting and sharing results | AR | R | I |
As the table shows, the security department holds accountability for plan formulation and detection/assessment, the AI/ML department is the final decision-maker for implementing improvements, and the business department is the final decision-maker for release approval. If a single work item is assigned A across multiple departments, it will cause stalling during emergencies over which department makes the final decision, so A must always be limited to a single department. The distinction between C (Consulted) and I (Informed) should also not be left ambiguous; including departments that only need to be informed within the "consulted" category will bloat meeting bodies and slow down the pace of the quarterly cycle.
Three Steps to Building an Internal Structure
Building the organizational structure proceeds with fewer do-overs if done in the order of current-state assessment, staffing, and process standardization. Even if roles and responsibilities are decided first, operations will break down if they diverge from existing operational capabilities, so the order of layering people and systems on top of the assessment results is the practical dividing line.
Step 1: Current State Assessment and Structure Design
In the current-state assessment, the department identifies the extent to which existing security operations can already address risks specific to LLMs. Even at companies that already have a general penetration testing framework in place, LLM-specific vulnerabilities such as prompt injection and hallucination are often not included in existing assessment items, and the starting point is work to visualize what is missing. Specifically, the following three points are organized as assessment items:
- The extent to which the existing security team has experience testing the categories in the OWASP LLM Top 10
- Whether the AI/ML department is in a state where it can disclose the contents of system prompts and fine-tuning to the security department
- Whether the business department has experience making decisions that connect vulnerability severity with business impact
If any of these three points is not yet in place, the organizational design must be structured to compensate for that gap. For example, if the security department lacks LLM-specific testing experience, a realistic design in the initial stage is to rely more heavily on the output of external LLM red-teaming frameworks and automated assessment tools, and to broaden the scope of human judgment only after in-house expertise has accumulated.
At the organizational design stage, rather than applying the RACI table as-is, the department holding A (Accountability) is provisionally adjusted according to the assessment results. In organizations where disclosure from the AI/ML department is insufficient, involvement from departments other than security is increased within the C (Consulted) role for test plan formulation, operating this as a transitional arrangement until the disclosure process is established. The results of this assessment and design are used as input information for the subsequent staffing and training programs.
Step 2: Staffing and Training Programs
Staffing begins by dividing dedicated and concurrent personnel according to the gaps identified in the current-state diagnosis. A practical arrangement is to assign one or two dedicated personnel from the security department to handle risks specific to LLMs, while the AI/ML department and business departments participate in the quarterly testing cycle on a concurrent basis alongside their existing duties. The reason for limiting dedicated personnel to specific departments is that knowledge of the OWASP LLM Top 10 categories and jailbreak verification methods is a cumulative skill, and this expertise tends to be lost if handled only on a concurrent basis.
The education program should be designed by separating classroom instruction from hands-on exercises. In classroom instruction, use the OWASP LLM Top 10 categories and the four functions of the NIST AI RMF (Govern, Map, Measure, Manage) as teaching materials, and have each person write out which of the company's business processes correspond to which function. In hands-on exercises, have participants use automated diagnostic tools such as Promptfoo or PyRIT to actually execute attack scenarios against existing test target systems, giving them experience in how to read detection results and how to distinguish false positives.
As a hypothetical example, having an AI/ML department staff member run an indirect injection scenario exercise against an internal inquiry system using RAG, to check how far instruction contamination via an external data source can get through the grounding check, functions effectively as educational content. After the exercise, recording which items each participant detected and which they missed, and adjusting the content of the next training session on an individual basis, helps reduce variation in proficiency levels.
Step 3: Standardizing the Testing Process
The purpose of standardizing the testing process is to create templates so that the judgments of personnel developed through staffing and the education program do not vary from person to person. The targets of standardization are three items: the format of the test plan document, the categorization of attack scenarios, and the recording items for detection results. If these are not documented and instead rely on verbal operation, know-how will not be passed on when personnel are transferred, and the quarterly cycle will end up depending on how long a given staff member remains in their position.
In the test plan document format, incorporate a checklist of the OWASP LLM Top 10 categories, and include a field clearly stating which categories are included in and excluded from the current test target. Recording the reasons for exclusion as well makes it easier to reconsider the decision when reviewing the test scope next time. For the recording items of detection results, use a minimum set of five items—attack pattern, success/failure, severity, responsible department, and scheduled retest date—and fix these as a template in a spreadsheet or ticket management system.
For cases such as excessive agency, where an AI agent executes operations on an external system, it is necessary to add a field for recording the side effects of execution results, unlike verification of a single prompt alone. For the categorization of attack scenarios, items with different verification procedures, such as direct injection and indirect injection, should be treated as separate categories rather than grouped into the same field, in order to prevent personnel from mistaking one procedure manual for another.
The standardized template should not be considered finished once created; it should actually be used in the education program's exercises, and items that are difficult to fill in or prone to omission should be reflected as feedback. After operations begin, both cases where the granularity of recording items is too fine and causes delays in entry, and cases where it is too coarse and prevents reproduction during retesting, should be subject to periodic review and adjusted in the revision for the next cycle. This template serves as the foundation for consistently operating everything from test planning to result evaluation.
Designing the Security Testing Process: From Planning to Reporting
To put the standardized template into actual testing operations, it is necessary to clearly specify who finalizes what at each stage of planning, execution, and reporting. In particular, the determination of the scope and the granularity of recording after detection are branching points where comparison during retesting becomes impossible if left to individual judgment.
Developing a Test Plan and Defining Scope
In drawing up the test plan, a key practical branching point is to define the scope not by "system unit" but by "function unit." Even within a single LLM application, the response generation portion of a customer-facing chatbot and the RAG search portion for internal inquiries differ in both attack vectors and verification procedures. Once the scope has been broken down into functional units, assign the OWASP LLM Top 10 categories to each function, and reflect in the schedule a separation between items that can be verified with a single input, such as direct injection, and items requiring multi-step verification via external data sources, such as indirect injection or RAG poisoning.
The handling of excluded items should also be finalized at the planning stage. If there are functions or attack categories not to be verified in the current cycle, the reason for exclusion should not simply be recorded as "low priority," but should instead be documented as a specific status, such as "verified in the previous cycle, no changes" or "carried over to the next cycle." Without this record, if the person responsible for drawing up the next test plan has been transferred, the background behind the exclusion becomes impossible to trace.
The estimation of the implementation period should be calculated individually, reflecting the breakdown by function unit and the differences in verification procedures. Estimating categories that only require single-step verification and categories that require multi-step verification with the same level of effort can cause only the indirect injection verification to fall behind during execution, which becomes a cause of the entire schedule collapsing. The exclusion field and period field in the plan document should receive final confirmation from the security department, which holds accountability according to the role-assignment table of the red team structure, after cross-checking consistency with the technical information provided by the AI/ML department.
Test Execution and Vulnerability Detection Flow
Test execution proceeds without duplication of work when the initial scan by automated diagnostic tools and the deep-dive verification by human judgment are kept separate. Automated diagnostic tools such as Promptfoo or PyRIT can run known jailbreak patterns and standard prompt injection attacks in bulk, so an automated scan should first be run against all target functions, and categories with high detection counts should be preferentially routed to human verification. Since the output of automated scans includes false positives, at this stage personnel should reproduce "whether an unauthorized output is actually being returned" one case at a time and exclude false positives.
In human verification, the focus should be on scenarios that are difficult for automated scans to cover comprehensively, such as attacks that go through external data sources like indirect injection, or jailbreaks that bypass guardrails through multi-turn dialogue. For functions using RAG, verification should involve actually injecting test documents to check how far malicious documents mixed into the vector database can get through the grounding check.
Detected vulnerabilities should be recorded on the spot in the five-item template fixed during test planning (attack pattern, success/failure, severity, responsible department, scheduled retest date), rather than operating in a way where entries are made all at once after verification is complete. Postponing the recording makes it more likely that duplicate records of the same pattern or omissions in entry will occur when multiple personnel are verifying in parallel. Finalizing the record each time verification of one function is completed, and strictly maintaining this order before moving on to verification of the next function, is a key practical point for preventing aggregation errors at the reporting stage.
FAQ: Common Challenges in Building an Internal Structure
Even when you understand the theory of building a governance structure, you may find yourself stuck on practical decisions such as the number of personnel, the temperature gap between departments, or how to judge the appropriateness of testing frequency. Below are answers to representative questions received during individual consultations.
How Can Small Companies Build a Red Team Structure?
For small companies, you should abandon the premise of placing dedicated personnel across multiple departments, and instead start with a scaled-down RACI in which a single security officer handles both planning and results evaluation, while in the AI/ML department, the development staff only concurrently handle improvement implementation. Approval from the business department can also be handled by a single member of management. Increase reliance on the output of automated diagnostic tools, and limit human verification to items that are difficult to automate, such as indirect injection or multi-turn jailbreaks. Splitting roles further as headcount increases is the approach least likely to cause the structure to break down.
What If Collaboration Between Security and AI Departments Isn't Working?
The cause is either insufficient information disclosure or a flaw in how discussions are designed. If the AI/ML department cannot disclose the system prompt or RAG configuration, treat this as a gap in the current diagnostic coverage. If disclosure is sufficient but meetings still stall, review whether departments unrelated to the "C" (Consulted) role in the RACI are being included, and whether consensus-building is being demanded for matters that could simply be handled as reports.
How Often Should Tests Be Conducted?
The basic cycle is roughly quarterly, but there are situations where the frequency should be moved up depending on trigger conditions.
- When system prompt or guardrail rules are changed: immediately retest only the scope affected by the change, within the relevant category
- When fine-tuning or a model update is performed: re-run testing across all categories without waiting for the quarterly cycle
- When the RAG data source or embedding model is changed: prioritize re-verification of RAG poisoning and grounding check items
- When a fix for a detected vulnerability is released: immediately retest only the relevant pattern, and reconfirm the whole scope during the quarterly cycle
Even during periods with no changes, continue the regular quarterly testing. Since attack methods and publicly disclosed jailbreak techniques keep evolving, resistance to new patterns can change even if the model or prompts remain unchanged. If you want to vary testing frequency by the importance of the feature, a practical approach is to run features with many users, such as customer-facing chatbots, on a cycle shorter than quarterly, while keeping low-frequency, internal-use features on the quarterly cycle.
ผู้เขียน・ผู้ตรวจสอบ
Yusuke Ishihara
เริ่มเขียนโปรแกรมตั้งแต่อายุ 13 ปี ด้วย MSX หลังจบการศึกษาจากมหาวิทยาลัย Musashi ได้ทำงานพัฒนาระบบขนาดใหญ่ รวมถึงระบบหลักของสายการบิน และโครงสร้าง Windows Server Hosting/VPS แห่งแรกของญี่ปุ่น ร่วมก่อตั้ง Site Engine Inc. ในปี 2008 ก่อตั้ง Unimon Inc. ในปี 2010 และ Enison Inc. ในปี 2025 นำทีมพัฒนาระบบธุรกิจ การประมวลผลภาษาธรรมชาติ และแพลตฟอร์ม ปัจจุบันมุ่งเน้นการพัฒนาผลิตภัณฑ์และการส่งเสริม AI/DX โดยใช้ generative AI และ Large Language Models (LLM)


