A failure that looked like a capability gap between coding agents was caused by how the API key was passed

日本語

Contents of this article

Summary

An AI agent compared three coding agents in a research repository that is scheduled to be made public. The research rebuilds an integrated agent development environment with OSS. The model and the tasks were the same for all three, and there were eight tasks. The source material consists of the records in this repository. In this comparison, a parent agent calls a parallel-execution MCP server. The parallel-execution MCP server is a program built as part of the research, and when it is called, it launches child agents in parallel. Only in the three runs with OpenHands CLI as the parent did none of the 32 child agents launched in total complete. Both the status field in the run record files and the result field of the measurement stayed normal, and for about a day this outcome was treated as a result of the comparison.

The cause was not a difference in capability between the coding agents but the way the measurement script was built. The measurement script left the delivery of the API key that the child agents use to two things: inheritance of the parent process’s environment variables and loading of a .env file. Only when the parent was OpenHands CLI and the run took place in an isolated working directory did the API key fail to arrive by either route.

The AI agent fixed the measurement script so that it writes the API key in the env field of the MCP server configuration and passes it that way. After the fix, the AI agent measured the condition with OpenHands CLI as the parent again, once each in two cases: with the set of instruction files in place and without it. The set of instruction files is a group of files, such as AGENTS.md, placed in the working directory. Under both conditions, the API key no longer caused every child agent to fail. With the set of instruction files in place, all eight child agents completed. Without it, 15 of the 16 child agents launched did not complete, for reasons other than the API key.

The first section of the body gives the overall picture of the cause and the fix. The sections after it are arranged in the following order.

The body names the products of the three coding agents compared. The content does not represent the views of any of their providers.

What you can take away

If you launch other agents from a coding agent through an MCP server, you will be able to do the following three things.

This article is a discussion based on the run records in the research repository scheduled to be made public and on the public source code checked at the time of writing.


Two implicit routes that the API key delivery depended on

This section first explains why every child agent failed in the comparison measurements and what the setup looks like after the fix.

The comparison measurements are a series of runs that measured three coding agents in turn, with the same model and the same tasks, while changing the conditions. A coding agent is a CLI program that is started from a terminal and that edits files and runs commands while calling a model. In the comparison measurements, the parent agent calls the parallel-execution MCP server. The parent agent is the coding agent process that receives the tasks and is started first.

The parallel-execution MCP server is an MCP server built as part of the research, and when it is called, it launches child agents in parallel. An MCP server is a program that provides tools for a coding agent to call through a common protocol called MCP. The MCP servers in this article are of the stdio type: the agent starts them as child processes and communicates with them over standard input and output. A child agent is a coding agent process launched by the parallel-execution MCP server.

The measurement script is a Python script that passes the tasks to the parent agent and starts it, then judges the results from the actual files and records them. Before the fix, the child agents’ API key arrived by one of the following two implicit routes. The API key is the credential for calling the inference service. The inference service is a service that makes models available to call through an API.

Only when the parent was OpenHands CLI and the run also took place in an isolated working directory did the API key fail to reach the child agents by either of the two routes. An isolated working directory is a separate working directory of the repository, created with git worktree. This summary combines three conditions found in the research records, and the research records contain no statement in this form. The three conditions are explained in the section “Isolating the cause by rerunning with the run records kept”.

After the fix, the measurement script writes the API key in the env field of the MCP server configuration and passes it explicitly. The env field of the MCP server configuration is a field in the coding agent’s configuration file. This field lists, for each MCP server, the environment variables to pass to the process that is started. The field is named environment in OpenCode and env in the other two.

When the API key is needed but missing, the parallel-execution MCP server stops before launching child agents and marks the run record file as failed. A run record file is a JSON file that the parallel-execution MCP server writes for each call. It has fields for the number of child agents launched, the number of child agents completed, and the status.

The following diagram shows the routes before and after the fix.

Before the fix
  The shell that started the measurement script (with the API key loaded)
    └ Parent agent
        └ Parallel-execution MCP server
            │  Route 1: environment variable inheritance
            │           Did not arrive when the parent was OpenHands CLI
            │  Route 2: read the .env file at the top level of the repository
            │           The isolated working directory has no .env file
            └ Child agent (the API key value is an empty string)

After the fix
  Measurement script
    └ Writes the API key in the env field of the MCP server configuration
        └ Parallel-execution MCP server
            ├ API key present: launches child agents
            └ API key needed but missing: stops before launching and marks the run record file as failed

Conditions of the comparison measurements and the number of completed child agents per parent agent

This section shows the conditions of the comparison measurements and the results for each parent agent in a table.

The conditions of the comparison measurements were as follows.

The following table summarizes, for each parent agent, the runs in which the parallel-execution MCP server was called.

Parent agent Runs in which a call occurred Run record files Child agents launched Child agents completed
Qwen Code as parent 3 runs 1, 1, and 2 per run 8 per run 8, 8, and 7 per run
OpenCode as parent 1 run 3 16 in total 8 in total
OpenHands CLI as parent 3 runs 4 in total 32 in total 0 in total

In the three runs with OpenHands CLI as the parent, 0 of the 32 child agents launched completed. In the run with OpenCode as the parent, the parent agent fixed an error in the orchestration script by itself and ran it again. The orchestration script is a script, run by the parallel-execution MCP server, that describes the steps for launching child agents.

Even in the three runs with OpenHands CLI as the parent, all eight tasks passed, and there were 0 rewrites of the test files. The records can be read as showing that the parent agent solved by itself the work that the child agents did not complete. However, the research records contain no breakdown of the parent agent’s work.

A judgment design in which the failure of every child agent did not show up in the results

This section explains how a result in which every child agent failed continued to be treated as a result of the comparison.

The measurement script judged whether the parallel-execution MCP server had been called by whether a run record file existed. Only the parallel-execution MCP server process writes run record files, so none is created if there is no call.

The measurement script recorded the numbers of child agents launched and completed by summing the fields of the run record files. For pass or fail on the tasks, it restored the test files, ran the tests for each task, and recorded the number of tasks that passed.

Even though every child agent failed, the failure did not show up in the results, for the following three reasons.

The result in which OpenHands CLI was the parent and 0 child agents completed was treated as a result of the comparison measurements for about a day.

Isolating the cause by rerunning with the run records kept

This section explains how the cause was isolated, in the order of the work. According to the research records, the AI agent did the implementation, ran the measurements, and handled the PRs. This AI agent works separately from the reviewing AI agent that examines the diffs of PRs.

The AI agent reran the condition with OpenHands CLI as the parent just once. This time, it kept the directory that holds the run record files instead of deleting it after the run. In the records it retrieved, all eight child agents had exited about 1.2 seconds after starting. The reason for exiting was an error saying that no authentication method had been selected.

The text of this error is in the file in Qwen Code version 0.21.13 that checks authentication in non-interactive mode. When Qwen Code starts in non-interactive mode, it determines the authentication method from the command arguments, the environment variables, and the configuration files. If none of them determines it, Qwen Code stops with this error.

The cause was the combination of the following three conditions.

  1. When the parent was OpenHands CLI, the parent process’s environment variables did not reach the parallel-execution MCP server process. When the parent was Qwen Code or OpenCode, they did reach it, so the API key loaded into the shell made it all the way to the child agents
  2. When the API key was not in the environment variables, the parallel-execution MCP server was built to read the .env file in the top-level directory of its own repository. The comparison measurements ran in an isolated working directory, and that directory had no .env file
  3. The measurement script did not write the API key in the env field of the MCP server configuration

When launching child agents, the parallel-execution MCP server put an empty string into the environment variable if it could not get the API key value.

Before that, there had also been a separate measurement with OpenHands CLI as the parent. In that measurement, the parent agent was explicitly made to call the parallel-execution MCP server, and one child agent had completed. The research records say that the explanation that fits best is that this run took place in the main working directory and the .env file was read. There is no record of which working directory it ran in, so this explanation has not been confirmed.

Differences between implementations in whether the parent process’s environment variables reach the MCP server

This section shows the results of checking, for each implementation in the public source code, whether the parent process’s environment variables reach a stdio MCP server. The checks were made on 2026-09-29.

The following table summarizes the environment variables that each implementation passes to the MCP server process when it starts a stdio MCP server.

Implementation Environment variables passed to the MCP server process Handling of the env field in the configuration Source
The official MCP Python SDK By default, only a list of environment variables that are safe to inherit. On POSIX, these are six: HOME, LOGNAME, PATH, SHELL, TERM, and USER Passed layered on top of the six Source code of the stdio client
The official MCP TypeScript SDK On POSIX, the same six as the Python SDK Passed layered on top of the list Source code of the stdio client
Qwen Code version 0.21.13 All of the parent process’s environment variables, excluding only three environment variables used to authenticate this product’s daemon and its child processes Passed layered on top of all of them Source code of the MCP client
OpenCode version 1.18.18 The whole of the parent process’s environment variables The environment field is passed layered on top of the whole Source code of the MCP handling
OpenHands CLI version 1.16.0 Reading the dependency path, the six defaults of the official MCP Python SDK The env field of the configuration is passed to the startup arguments of the official MCP Python SDK Dependency definition, MCP handling in openhands-sdk, stdio handling in fastmcp

Notes on each row are as follows.

The OpenHands CLI row is an explanation obtained by reading the source code of the MCP part. This explanation fits best with the observation that the environment variables did not arrive when the parent was OpenHands CLI. However, the exact versions of fastmcp and the official MCP Python SDK installed at the time of the measurements have not been checked. So this explanation is not the result of checking the whole route in the measurement environment. The observed fact is limited to the way these measurements were started. When the parent was OpenHands CLI version 1.16.0, the parent process’s environment variables did not reach the MCP server.

The official documentation of the three coding agents describes the field for passing environment variables. The pages checked are the following three.

On the other hand, none of the three pages says whether a started MCP server receives the whole of the parent’s environment. The inheritance behavior is determined only by the implementation in the source code, and it can change between versions.

Both designs have reasons behind them. The default in the official SDKs narrows the inherited environment variables to avoid leaking secret values. Qwen Code limits what it excludes to values for the product’s internal use. The comments in the source code give the reason that commands users run in the shell depend on inheriting third-party credentials. Qwen Code version 0.21.13 passes the parent’s environment variables only through that same function when it starts an MCP server as well. The measurement script after the fix does not depend on either design. This is because the measurement script writes the API key in the env field of the MCP server configuration and passes it explicitly.

The fix that passes the API key explicitly in the env field of the MCP server configuration, and the re-measurement

This section shows the fix to how the API key is passed and the results of the re-measurement after the fix.

The AI agent fixed the measurement script. After the fix, the measurement script writes the API key that the child agents use in the env field of the MCP server configuration and passes it explicitly. The AI agent also fixed, in the same way, a separate measurement script that explicitly makes the parent agent call the parallel-execution MCP server. That values written in the env field reach the MCP server in all three coding agents had already been confirmed earlier, by passing a different value.

The AI agent set up a unit test to confirm that the API key is written in OpenHands CLI’s mcp.json. After that, it extended the unit test to all three coding agents.

The following examples show the form for writing the API key in the env field of the MCP server configuration, written from scratch to match the form in the official documentation. The server name, the environment variable name, and the value are placeholders. The first is the Qwen Code form, and OpenHands CLI’s mcp.json uses the same form. The second is the OpenCode form.

{
  "mcpServers": {
    "parallel-runner": {
      "command": "python",
      "args": ["-m", "example_parallel_runner"],
      "env": {
        "EXAMPLE_API_KEY": "replace-with-your-api-key"
      }
    }
  }
}
{
  "$schema": "https://opencode.ai/config.json",
  "mcp": {
    "parallel-runner": {
      "type": "local",
      "command": ["python", "-m", "example_parallel_runner"],
      "environment": {
        "EXAMPLE_API_KEY": "replace-with-your-api-key"
      }
    }
  }
}

After the fix, the AI agent measured two conditions with OpenHands CLI as the parent again, once each. The two conditions differ in whether the set of instruction files for agents is in place. The set of instruction files for agents is the group of AGENTS.md, skill files, and memory files placed in the working directory. The following table shows the results of the re-measurement under the two conditions.

Condition Child agents launched Child agents completed Tasks passed Time taken
Set of instruction files in place 8 8 8 of 8 556.1 seconds
Set of instruction files not in place 16 in total 1 in total 8 of 8 761.1 seconds

The time limit was 900 seconds. Under the condition with the set of instruction files in place, the parent agent called the parallel-execution MCP server on its own. All eight child agents completed. This was the first run, under a condition with OpenHands CLI as the parent, in which the parent agent called the parallel-execution MCP server on its own and every child agent launched completed.

Under the condition without the set of instruction files, a call also occurred, and there were three run record files. Of the 16 child agents launched under this condition, 15 did not complete, for reasons other than the API key. The research records say that these 15 include the pattern of fixing an error in the orchestration script and running it again. The research records also say that observation of the causes of these 15 will continue.

From these results, the AI agent concluded as follows. The cause of 0 completed child agents was not a difference in capability between the coding agents but the way the measurement script was built. In response to feedback from the reviewing AI agent, the AI agent rewrote the scope of the conclusion. The reviewing AI agent is an AI agent that is started after a PR is created, separately from the agent that did the implementation. It examines the diff read-only. In the rewritten conclusion, the AI agent stated definitively that the problem of every child agent failing because of the API key had been resolved. It separated out the remaining 15 incomplete child agents as having reasons other than the API key.

The two re-measurement runs were recorded separately from the comparison totals, as additional measurements for isolating the cause. The re-measurement was one run for each condition. So these records cannot tell how often every child agent launched completes when the parent is OpenHands CLI.

A mechanism that does not hide failures, and unit tests that prevent the same omission

This section explains the mechanism that does not hide failures and the unit tests that prevent the same omission, both added in PRs after the API key fix.

The AI agent added handling for when the API key is needed but missing. In this case, the parallel-execution MCP server stops before launching child agents and marks the run record file as failed. It does not treat the stop as the failure of one child agent, and marks the whole call as failed.

The PR that added this handling went through three reviews.

  1. In the first review, the reviewing AI agent demonstrated a problem by actually running the code. The problem was that the exception from the stopping logic was treated as the failure of one child agent, and the status in the run record file was left as normal
  2. In the second review, the reviewing AI agent showed that with the parallel and sequential ways of writing the orchestration, the same exception was still treated as normal
  3. In response to the third review, the AI agent changed the stopping logic so that, once it runs, it leaves a failure flag in the run record file. This flag remains even if the orchestration script catches the exception and swallows it

In the same PR, the AI agent removed the logic that, when the API key was empty, overwrote the caller’s environment variable with an empty string. After the fix, when there is no API key, the parallel-execution MCP server passes the authentication in the caller’s environment to the child agents as it is.

In the same PR, the AI agent also added logic that automatically puts a flag on the rows the measurement script records, indicating that there was a defect on the measurement side. Rows with this flag are not counted in the totals as observations of an agent’s capability. The flag is added when a task run meets all three of the following conditions.

The conditions were made narrow because, if failures caused by a lack of capability were excluded from the totals as defects on the measurement side, the totals would look better than the actual capability. The logic that adds this flag is separate from the logic that judges whether the parallel-execution MCP server was called.

The same kind of omission also happened with another environment variable. An isolated directory is passed through an environment variable to Qwen Code, the child agent. This directory is used in place of the user’s home directory. The measurement script set this value at the point of starting the parent only when the parent was Qwen Code. So this value reached the child agents through environment variable inheritance only when the parent was Qwen Code. When the parent was one of the other two, the value did not exist at the point of starting the parent, so it did not reach the child agents.

The AI agent changed this value as well so that it is written in the env field of the MCP server configuration and passed explicitly for all three coding agents. It confirmed this setup with a unit test. The reviewing AI agent pointed out two places with the same omission. The two places were a separate measurement script and an integration test. The AI agent fixed both places in the same way.

In a later PR, the AI agent consolidated the logic that builds the env field of the MCP server configuration into a single function. This is because, if callers set values separately, the same omission happens when the number of places to register grows. The unit test confirms that the set of environment variable names in the env field exactly matches a defined set. If a new credential is written into a configuration file without anyone noticing, this unit test fails.

Protection added because the API key is now written in the env field

This section explains the new risk created by the fix that passes the API key explicitly, and the protection added in the same PR.

After the fix, the API key written in the env field is written as a raw value into a configuration file inside the working directory. The reviewing AI agent returned seven findings on this PR. Each finding is given one of four severity levels: Critical, High, Medium, or Low. All seven had Medium severity. In response to the findings among the seven that concerned protecting the API key, the AI agent fixed the following two things.

Two rules for designing comparison measurements

This section presents two rules for designing comparison measurements that can be generalized from these records. The research records contain no statement in this form.

  1. Before starting comparison measurements, identify the implicit behaviors that differ between coding agents. Whether the parent process’s environment variables reach the MCP server is one of those behaviors
  2. Do not leave credentials to environment variable inheritance, and pass them explicitly from the measurement script

The basis for rule 1 is the observation that only when the parent was OpenHands CLI did the environment variables fail to arrive, and the source code in which the inherited range differs by implementation. The basis for rule 2 is the facts written in the section on the fix that passes the key explicitly in the env field and the re-measurement, and in the section on the mechanism that does not hide failures.

As a supplementary note, the same kind of failure also happened on an external evaluation execution platform. The AI agent ran Qwen Code over ACP on the external evaluation execution platform. ACP is a public protocol for connecting editors and coding agents. This run got as far as starting Qwen Code and connecting over ACP, and then failed with an error at the authentication stage. The task score was 0. In the logic that started the execution platform at that time, the AI agent could not find a configuration item for passing the model’s authentication information to the container. The research records treat this as the cause. The measurement script in the same repository had gotten through the same authentication by passing the environment variables and a copy of the configuration file. The AI agent recorded this run as a defect on the measurement side, not as an observation of capability.

The inheritance behavior of the parent process’s environment variables can be checked in the source code of the version in use, not in the official documentation. The part to read is the logic that starts a stdio MCP server. It is a good idea to check whether that logic passes the whole of the parent’s environment or narrows it with a list, and only then use that version of the coding agent in comparison measurements.

Materials I referred to

The public materials referred to in the body are as follows. All of them were checked on 2026-09-29.