What Are Strix, Cairn, and Hermes? How AI Security Tools Work and Their Risks

What Are Strix, Cairn, and Hermes? How AI Security Tools Work and Their Risks

Strix, Cairn, and Hermes are tools that use AI for vulnerability research, execution, and ongoing task management. Rather than viewing them as competing products with the same functionality, understanding them as AI agents with different roles reveals their distinctions.

Some readers may have first learned of these three tools through coverage by Weekly ASCII. What businesses need to focus on is not the novelty of the names, but rather "under what authority can AI-determined actions be executed?" This article digs into the design of each tool, explains how to interpret the numbers behind reported misuse cases, and outlines the decision criteria businesses should apply when adopting these tools for defense or verification.

How Do the Roles of Strix, Cairn, and Hermes Differ?

To start, these tools can be organized by three roles: exploration, goal-directed execution, and ongoing management. That said, their actual functionalities overlap in practice.

Separating Roles in the Incident from Product Features

In Gambit Security's investigation, Strix was used for vulnerability exploration, Cairn for intrusion work, and Hermes for job initiation and overall activity management. This division of labor reflects the specific incident in question, not a classification that limits each product's functionality. Gambit's primary report

Below is a breakdown of functionality based on publicly available materials. The row for Cairn refers to the public version, oritera/Cairn, and it has not been confirmed that this is identical to the implementation used in the incident.

ToolPerspective on roleWhat to verify during evaluation
StrixInvestigates suspicious areas and works toward proving the issueWhether evidence can be reproduced and leads to a fix
Cairn (public version oritera/Cairn)Selects the next exploration step based on the gap between current state and goalWhether completion criteria and stopping conditions are clear
Hermes AgentUses memory and skills to continue and manage ongoing workWhether it can manage permissions, memory, and scheduled execution

The evaluation items in the table serve as decision criteria for adoption. They are not a performance ranking based on measuring all three products under identical conditions. The outcomes that matter will differ depending on whether you're the one receiving diagnostic results, overseeing verification, or managing operations.

The AI Model and the Mechanism That Operates Tools Are Different Things

Even when an AI model determines the next action from text, that alone doesn't make a browser or device actually move. Software is needed to connect the model's output to tool invocations, return results, and continue the next round of decision-making. This supporting mechanism for execution is also referred to as a harness. Strix's official documentation describes a design that combines a browser, a proxy for inspecting HTTP communications, a terminal, and other components with the agent. Strix official documentation

What this tells us is that risk is not determined solely by the model's intelligence on its own. Even if the same judgment is produced, the impact differs between an environment where the only permission is to generate a report and one where the system itself can be modified.

In internal evaluations as well, in addition to "which model is being used," organizations should check what operations can be executed, what assets can be accessed, how decisions are logged, and under what conditions approval is required. Evaluating only whether a conversational response is accurate is insufficient for managing an agent that performs actual execution.

In fact, in the reported case, the three tools each used different AI models via provider services. It is particularly notable that when the managing tool had its requests refused by a newer model, the attacker switched to an older-generation model. This suggests that while safety measures on the model side can serve as a certain deterrent, room remains to escape to older models or alternative routes.

For more on the concept of dividing work among agents, the explanation of agent orchestration may also be helpful.

How Does Strix Investigate Vulnerabilities?

A defining feature of Strix is that it attempts to connect the discovery of suspicious areas to verification of whether an actual issue holds up. It is released as a penetration testing tool aimed at developers and security personnel.

Verifying by Combining Code, Communications, and Screens

Strix is equipped with the capability to examine the behavior of running applications in addition to code-reading analysis. According to official materials, the architecture is described as one in which specialized agents collaborate to share findings and confirm vulnerabilities through demonstrations called PoCs. Strix Official Documentation

For example, suppose that in an internal test application, a finding points out that "information might be visible to people other than the intended personnel." Relying only on conditional branches in the code leaves open the possibility that actual access control operates at a different layer. Conversely, even if a button disappears from the screen, that does not necessarily mean that permission checks on the server side are also correctly implemented. While this is merely an illustrative example of the design, it conveys the significance of cross-checking multiple perspectives.

However, diagnostics that involve actually running and verifying the system can entail communication with the target and changes in its state. This cannot be treated the same as a review that simply reads the source code. The person in charge needs to receive results in a form that allows them to trace what operations were necessary to confirm what.

Even With Proof-Backed Findings, There's Room for Human Judgment

Strix's official repository lists, as its features, findings accompanied by proof, remediation suggestions, report generation, and integration into the development workflow. However, such product descriptions are not an independent guarantee that there will be no oversights or false positives in all environments. Strix Official Repository

At the time of acceptance, one should cross-check at least the version of the target, the account privileges used, the expected behavior versus the observed results, and the conditions for reproduction. It is also important not to treat a phenomenon that occurs only under strong test-only privileges as directly representing the possibility of intrusion from outside. Conversely, one cannot conclude that something is "safe" merely because verification failed.

The same applies to proposed fixes. Even if a fix blocks unauthorized access, if it also causes legitimate users' operations to fail, it cannot be adopted as is. Rather than increasing the number of detections, there is value in narrowing down to important issues with supporting evidence and verifying business operations after the fix as well.

How to evaluate AI's diagnostic capability is also discussed in An Explanation of CyberGym and AI Security Evaluation.

How Does Cairn Advance Its Goal-Directed Exploration?

Cairn serves as a clue for understanding the mechanism by which the next exploration is chosen based on the information obtained. However, Gambit's report does not specify the version used or the repository, so it cannot be definitively concluded that the implementation involved in the incident is identical to the open-source project oritera/Cairn discussed here. Gambit's Primary Report

Sharing Verification Results and Exploration Plans to Choose the Next Action

The public repository oritera/Cairn positions itself as a general-purpose state-space exploration engine, with penetration testing as its first area of validation. At its core is a shared bulletin board that handles Facts (confirmed results), Intents (unexecuted exploration strategies), and Hints (judgment material provided by humans). Workers do not have fixed roles; they read the shared state, explore, and write results back. Cairn Public Repository

This design differs from an approach in which a long checklist is created once and then worked through in order. This is because, if the premises change midway through the work, what needs to be checked next also changes.

For example, in an authorized internal test, if it turns out that a certain feature has already been deprecated, the rationale for continuing to investigate that feature weakens. If the confirmed results can be shared, other workers can use the same premise. This is an example intended to illustrate the mechanism, not a demonstration of actual diagnostic results.

Shared Facts Must Also Be Externally Verifiable

In Cairn's design documentation, the initial attempt, the judgment to re-read the situation, and individual exploration are treated as separate tasks. The dispatcher is responsible for task assignment and execution management, while workers handle situational judgment and individual exploration. Cairn's Dispatcher Design

Given this structure, the quality of shared information affects subsequent judgments. Even if something is stored under the name "Fact," if an external evaluator cannot verify that its content is correct, there is a risk of spreading a mistaken premise across multiple tasks. Observed evidence, the target, the time, and the interpretation should be recorded separately, and even a report stating that "the goal was achieved" should be required to have verifiable grounds.

The public version of Cairn explicitly states that there is no sandbox for local execution. How to restrict the privileges of the execution environment is an important item in deployment decisions. Cairn Public Repository

What Does Hermes Remember, and How Does It Manage Tasks?

Hermes Agent is a general-purpose agent published by Nous Research. The fact that it has memory and skills, and can continue working beyond a single conversation, is key to understanding its role as a manager. Hermes official repository

Memory, Past Conversations, and Skills Each Play a Different Role

Hermes's persistent memory is a mechanism for storing key points to be used across sessions. According to the official documentation, memories such as information about the environment are separated from information about the user, and are loaded at the start of a session. Detailed past context is supplemented by searching session history. Hermes's persistent memory

Skills are documents containing procedures and knowledge that are loaded when needed. Rather than loading all content every time, the structure proceeds from a name or summary to the necessary document. Hermes's skills specification

In terms of business work, it is easier to understand if you think of memory as the key points of a handover, history as a record of past work, and skills as reusable procedure manuals. The description of "learning from experience" is also best understood by starting from this storage and reuse. These features alone are not grounds to conclude that the weights of the AI model itself are retrained each time.

Because It Can Continue, Old Decisions and Permissions Must Also Be Managed

Hermes's official repository also documents features such as scheduled execution and delegation to sub-agents. These features can also be used for regular investigations and repetitive tasks. It is inappropriate to regard Hermes itself as a product dedicated solely to attacks. Hermes official repository

In operation, what matters is what is retained in the continuing mechanism. Even if, hypothetically, a target or operation that was permitted during last month's verification was saved in a skill, that permission is not necessarily still valid this month. Whether a procedure is useful and whether its execution is approved are separate issues.

Possible countermeasures include keeping a change history of procedure manuals, checking the expiration of permissions at the time of execution, and designating a person responsible for scheduled jobs. At the end of a task, it is not enough to simply close the conversation; scheduled processes and work delegated to other agents should also be checked. The memory and automated execution that support convenience are only useful for business operations if they can be properly updated and stopped.

Hermes's official documentation describes a mechanism that requires approval before executing dangerous commands, as well as a setting that skips the approval confirmation. However, some restrictions remain, such as an always-active prohibited list. When adopting the system, be sure to check both the approval settings and the restrictions that remain. Hermes's security documentation

What Can We Learn From the Reported Attack Case?

The important point about this case is that the burden of continuing exploration and execution can vary depending on the AI. Figures for damage and cost need to be read with attention to matching the subject of investigation and the aggregation conditions. According to Gambit's report, 105 attack projects were launched between September 10 and 15, 2026, and at least 27 companies were compromised in some form.

The Average of $25.46 Is Not the Total Cost per Breach

The average of $25.46 recorded by Gambit is the attackers' own cost tally, representing the average amount per scan across 101 completed scans. There is a range, from $3.13 for the cheapest target to $79.31 for the most expensive. This is a tally centered on model usage fees, and is neither the total cost of a single successful intrusion nor the price of a commercial service. In a separate tally, $7,005.71 in model usage fees over four weeks was confirmed from account records, and the cost including subsequent activity is estimated at roughly $12,000 to $18,000. Gambit's primary report

This tally cannot be directly linked to the success rate of intrusions or the amount of damage to companies. Even when considering adoption at your own company, this figure should not be used as-is as a budget guideline. Besides model usage fees, there are burdens involved in preparing the environment, checking results, making corrections, and re-verifying.

In evaluation, it becomes easier to grasp the actual situation by separating "how much it cost to run once" from "how much it cost to obtain results usable for decision-making." Even if a large number of low-cost findings are produced, if human verification takes a long time, the overall workload may not necessarily decrease.

Proceeding With Minimal Instructions Is Different From Full Autonomy

The report covers activities that include human instructions as well. For some targets, humans provided valid administrator passwords to the agent, so not all intrusion paths were opened by the AI on its own. The investigation's evidence base also includes AI-generated reports, and the investigators themselves note the possibility of errors. Gambit's investigation methodology

Therefore, this case alone cannot be generalized to mean that "anyone can breach any company just by giving a short instruction." On the other hand, given that there is a mechanism that reads intermediate results, selects a different approach, and continues working, defenders cannot assume that the attacker will give up after a single failure. This is an operational implication that can be drawn from the design of each tool and from the case itself.

The focus of countermeasures should also not be limited to simply blocking a single suspicious communication. Logs that allow investigation of how multiple operations connect across time, and whether permission changes or new scheduled processes have been created, become important.

Where Should Companies Start Preparing?

Preparedness is needed both for countermeasures against external breaches and for permission management of agents operated in-house. When starting your own testing, decide on objectives and stop conditions first, then evaluate on a small scale.

Separately Inspecting External-Facing Areas, Authentication, Payment Screens, and Recovery

It is important that a defense review not stop at just the entry point. In addition to externally exposed systems, administrator privileges, distributed files, and the recovery means for resuming business should all be checked, with responsibilities divided among different owners.

Target to checkQuestion the owner should be able to answer
Externally exposed systemsIs the owner and update status known, and can fixes be advanced according to exploitation risk?
Administrator / service privilegesAre unnecessary privileges removed, and can usage be tracked?
Payment page scriptsCan approval and integrity be confirmed, and can tampering with the page delivered to users and the related HTTP headers be noticed?
Backup / recoveryIs there a copy that is protected even if production privileges are compromised, and can it be restored?

For payment pages, monitor not only the files on the server but also the page actually delivered to the user's browser and the HTTP headers that affect security. Even if tag management is handled by a separate department, confirm the whole picture and decide in advance who is responsible for stopping things in the event of an anomaly.

The reported case shows why this preparedness for recovery matters. The major data loss was not caused by a ransom demand but arose from the attacker's cleanup process. A procedure to delete data after extraction had been planted, and at one retailer, the matching for deletion targets by name was too broad, resulting in the deletion of 180 tables, including the administrator's backup. It is necessary to secure a copy that cannot be altered or deleted even with compromised production privileges, and to confirm that it can be restored.

For prioritizing fixes, CISA's KEV Catalog, which collects vulnerabilities confirmed to have been actually exploited, can serve as a reference. For countermeasures against payment page tampering, PCI SSC's explanation of script management and monitoring provides the basis, and for recovery preparedness, CISA's guidance on backup and restoration verification provides the basis.

If you want to broadly organize preparedness against AI-driven attacks, also see the article on AI cyberattacks and corporate defense.

Combining Instructions to the Agent With Actual Permission Restrictions

OWASP organizes the problem of giving LLM agents excessive functionality, permissions, and autonomy as Excessive Agency. It recommends limiting agents to the necessary tools and permissions, and incorporating human approval for high-impact operations. OWASP's explanation of Excessive Agency

Permissions for diagnostics should also be considered separately: the permission to save reports and the permission to rewrite targets. In addition to instructions like "do not make changes," a configuration that does not grant unnecessary change permissions is needed.

There is also the risk of "indirect prompt injection," in which malicious instructions embedded in the pages or files being diagnosed induce the agent to act. This is not a technique confirmed in this particular incident, but a general risk in deployment. External content should not be treated as execution instructions, access to sensitive information and external transmission should be restricted, and approval should be required for high-impact operations. OWASP's explanation of prompt injection

Hermes's official security documentation also explicitly states that its file operation protection features are not a sandbox that isolates malicious or compromised agents, since terminal operations can be a separate pathway. Hermes's security documentation

When separating execution environments, also inspect shared folders, external communications, and credentials brought into the environment. If strong permissions are handed over inside an isolated environment, the impact on whatever that permission can reach still remains.

Using Small-Scale Tests to Measure Verification and Fix Burden, Not Just Findings Count

If using this for defensive purposes, first prepare a limited test environment with written authorization. Strix's official documentation also requires testing only applications you own or targets for which you have explicit permission. Strix official documentation

Before testing, decide on the target, accounts, execution time, the scope of changes allowed, and the stop conditions in case of anomalies. Also align the versions, settings, and evaluation targets being compared.

Even if you run the tool locally, in configurations that use a cloud model, data is sent to the provider. Check the scope of transmission and masking for code, communication content, and credentials during diagnostics, as well as the provider's storage and training-use terms. Strix's explanation of model execution environments

Metrics to measure include the proportion of findings that could be reproduced, the time required for human verification, missed critical issues, and the time required for re-verification after fixes. This is a recommended evaluation method, not actual measured values for the three tools. Preparing a training application containing known issues, along with a version where the issues have been fixed, allows detection and fix-verification to be considered separately.

Checking the stop operation should also be included in the evaluation. Confirming whether terminal processes, sub-agents, and scheduled jobs all stop when the person in charge instructs an interruption helps clarify the scope of responsibility after deployment.

Frequently Asked Questions About Strix, Cairn, and Hermes

When considering adoption, you naturally become concerned about cost and how it relates to human-performed diagnostics. You should separate the question of the software's release format from the burden of using it in actual operations.

If a Tool Is Publicly Available, Is It Free to Use?

Using publicly released software is not the same as incurring no operating costs. In addition to the open-source version that runs locally, Strix offers other release formats such as a cloud version. The terms of use for the model, computing resources, and the costs of maintenance and verifying results need to be considered separately. Strix's release formats

The three publicly released pieces of software also differ in license. Strix uses Apache-2.0, Hermes uses MIT, and the public version of oritera/Cairn uses AGPL-3.0. Cairn's official README states that if you wish to use it in a commercial or proprietary environment without being bound by AGPL obligations, you should obtain a separate commercial license. Cairn's license description

Before adopting any of these, be sure to check the terms in each repository's license.

When comparing them, first align the version you intend to use with the contract terms. The costs reported in connection with this incident cannot simply be substituted for the cost of operating the software safely within your own organization.

Will Human Vulnerability Assessments and Security Staff Become Unnecessary?

Based solely on the materials for these three tools and this incident, it cannot be concluded that human diagnosis is no longer necessary. Even if these tools can assist with investigation and verification, agreement on the scope to be covered, judgment of the impact on operations, confirmation of the validity of results, and prioritization of fixes remain necessary.

A practical approach is to first assign roles in situations where outcomes are easy to cross-check, such as recurring periodic verification or re-confirmation after a fix. Rather than focusing on how many staff members can be reduced, it becomes clearer what adoption actually means if you evaluate how far you can now verify areas that previously could not be fully checked.

The division of responsibility and judgment between personnel is explained in detail in Building an Internal Red Team Organization.

Summary: Evaluating Exploration, Execution, and Continuous Management Through Evidence and Permissions

According to publicly available materials, Strix is characterized by investigation and verification, oritera/Cairn by goal-driven exploration, and Hermes by ongoing management through memory and skills. These roles overlap to some extent, and the version and configuration used in the incident must be considered separately from the currently published features. In particular, for Cairn, it has not been confirmed whether the implementation used in the incident is identical to the publicly released version.

For a company, the criteria for judgment are not limited to how autonomously the AI operates. Only by also examining whether there is evidence to verify results, whether only the necessary permissions have been granted, whether the process can be halted midway, and whether business operations can be restored, can you properly judge whether to adopt it.

When considering use within your own organization, first determine the authorized testing environment and the conditions under which results will be accepted. When consulting about AI adoption or reviewing business systems, organizing information on the target system, current permission management, and the issues you want to address will lead to more concrete discussions.

Author & Supervisor

Yusuke Ishihara

Yusuke Ishihara

Started programming at age 13 with MSX. After graduating from Musashi University, worked on large-scale system development including airline core systems and Japan's first Windows server hosting/VPS infrastructure. Co-founded Site Engine Inc. in 2008. Founded Unimon Inc. in 2010 and Enison Inc. in 2025, leading development of business systems, NLP, and platform solutions. Currently focuses on product development and AI/DX initiatives leveraging generative AI and large language models (LLMs).