> ## Documentation Index
> Fetch the complete documentation index at: https://aspect.build/llms.txt
> Use this file to discover all available pages before exploring further.

# You won't ship that

> Attach the same analysis to the commands you already run, using trait flags and lifecycle hooks in config.axl.

```shell theme={null}
git checkout step-6
git diff step-5 step-6
```

`aspect observe` taught you what's in the streams. But nobody is going to type `aspect observe` instead of `aspect test`, in CI or anywhere else. A command that competes with the one people already use loses.

You don't rewrite `aspect test`. You attach to it.

## Seventeen lines

```python theme={null}
load("@aspect//traits.axl", "BazelTrait")

def config(ctx: ConfigContext):
    ctx.traits[BazelTrait].extra_flags = ["--flaky_test_attempts=3"]

    caught = []

    def _watch(ctx, event):
        if event.kind == "test_summary" and event.payload.overall_status == "flaky":
            caught.append(event.id.label)

    def _alert(ctx, conclusion):
        if caught:
            ctx.std.io.stderr.write("\n  ⚠  FLAKE CAUGHT: " + ", ".join(caught) + "\n")

    ctx.traits[BazelTrait].build_event.append(_watch)
    ctx.hooks.post_task(_alert)
```

That is the build-event half of `.aspect/config.axl`. The file on `step-6` is 32 lines, because it already carries the execution-log half too — that is the section after next, and the two are shown apart only because they are two different doors. `_alert` writes to your terminal, but it is ordinary Starlark in an ordinary function; a webhook POST would go in the same place.

Now run the tests you were going to run anyway. One addition: step 1 already ran them, so without `--nocache_test_results` Bazel hands back 55 cached passes and no test executes, which means no flake can fire.

```shell theme={null}
aspect test //... --nocache_test_results
```

Bazel's progress streams exactly as before. The flaky test fails on close to one run in three, so this is a dice roll. Sometimes all three attempts fail and you get a plain `exit code 3` with no `FLAKE CAUGHT` line at all — the retries are correlated, so that is commoner than one-in-twenty-seven. Run it again. When it comes up flaky, you get this at the end:

```text theme={null}
  ⚠  FLAKE CAUGHT: //internal/domain/forecast:forecast_test
  ⚙  EXECUTED: TestRunner×56

→ ⚠️  Flagged test task in 798ms · Tests flaky
```

Nothing replaced, nothing hidden, nothing new to remember.

Fifty-six, not fifty-five. A clean run executes 55 `TestRunner` spawns, one per test. The retry Bazel ran to prove the flake is a 56th spawn, and it is not a cache hit, so `_spawn` counts it. Three attempts gives 57. The two halves are reading the same build from different ends, and the count is the proof.

## The two halves chain

This is worth slowing down for, because the order is not obvious.

**Without `--flaky_test_attempts`, a flake is indistinguishable from a failure.** Bazel runs the test once, it fails, you get `failed`. Only when Bazel retries and the retry passes does it report `flaky` — so the trait flag is what makes the thing *observable* before the hook can observe it.

Prove the flag reaches a command you never touched:

```shell theme={null}
aspect test //lib/queue:queue_test --announce-bazel-command=true 2>&1 | grep -o -- '--flaky_test_attempts=[0-9]*'
```

```text theme={null}
--flaky_test_attempts=3
```

And confirm whose command that is:

```shell theme={null}
aspect describe test | jq '{command, defined_in}'
```

```json theme={null}
{
  "command": "aspect test",
  "defined_in": "@aspect//test.axl"
}
```

`@aspect//test.axl` — a file you have never opened, carrying a flag from your `config.axl`.

## What `caught` is doing

The two hooks share a list through a closure. `_watch` runs per event, during the build; `_alert` runs once, after the task body. Neither knows about the other, and nothing plumbs state between them.

That is the thing worth taking away. Expressing "accumulate during, report after" in a configuration *schema* means inventing somewhere to put the accumulator. In a language it's a variable.

## What this does and doesn't reach

`build_event` fires for **every** task that drives Bazel — `test`, `build`, `lint`, `format`, `gazelle`, `delivery`. One hook, every command.

The execution log works the same way, and it is the half worth coming back for.

<Note>
  `--flaky_test_attempts` changes what the build **reports**, not what it does. The test is still flaky; you have only made it visible. Shipping that flag to CI because it turned a red build green is how a flake becomes permanent.
</Note>

## The other stream, without a task of your own

Step 5 read the execution log live, out of a task that spawned Bazel itself. That second part is what does not travel: `build.execution_logs()` needs a `build` *you* spawned, and `config.axl` spawns nothing — it configures the invocation somebody else's `aspect test` is about to make. So the trait carries its own door, the same shape as `build_event`: a per-entry callback.

```python theme={null}
load("@aspect//traits.axl", "BazelTrait", "ExecLogHook")

def config(ctx: ConfigContext):
    ctx.traits[BazelTrait].extra_flags = ["--flaky_test_attempts=3"]

    caught = []
    executed = {}

    def _watch(ctx, event):
        if event.kind == "test_summary" and event.payload.overall_status == "flaky":
            caught.append(event.id.label)

    def _spawn(ctx, entry):
        spawn = entry.type
        if not spawn.cache_hit:
            executed[spawn.mnemonic] = executed.get(spawn.mnemonic, 0) + 1

    def _alert(ctx, conclusion):
        if caught:
            ctx.std.io.stderr.write("\n  ⚠  FLAKE CAUGHT: " + ", ".join(caught) + "\n")
        if executed:
            top = sorted(executed.items(), key = lambda kv: -kv[1])[:3]
            ctx.std.io.stderr.write(
                "  ⚙  EXECUTED: " +
                ", ".join([k + "×" + str(n) for k, n in top]) + "\n",
            )

    ctx.traits[BazelTrait].build_event.append(_watch)
    ctx.traits[BazelTrait].exec_log_event.append(
        ExecLogHook(on_entry = _spawn, kinds = ["spawn"]),
    )
    ctx.hooks.post_task(_alert)
```

`_spawn` reads an entry the same way step 5's loop did — `entry.type` is the payload, `type(entry.type)` is the kind, and a spawn's fields are the same ones. It asks different questions of it: step 5 keyed on `spawn.runner` for the local/remote/cache breakdown and on `spawn.input_set_id` to reconstruct input sets, this one keys on `spawn.cache_hit` and `spawn.mnemonic`. What is gone is the drain loop, and the invocation you had to own in order to have one.

Run the tests you were going to run anyway, again:

```shell theme={null}
aspect test //... --nocache_test_results
```

```text theme={null}
  ⚙  EXECUTED: TestRunner×55
```

Fifty-five executed `TestRunner` spawns, which is exactly what Bazel said about itself on that run — `55 darwin-sandbox`. Everything else came back a cache hit, and `--nocache_test_results` is the only reason the tests did not. On a tree you have already built, that is the whole list. From genuinely cold it is longer — `GoCompilePkg×222, GoLink×62, GoTestGenTest×55` — and third place is a three-way tie at 55 between `GoTestGenTest`, `GoCompilePkgExternal` and `TestRunner`, so the third entry moves between runs. Read the count, not the ordering. `_spawn` ran while Bazel was executing, on entries it was handed as they were produced, and it saw every one of them: a hook is as live and as complete as step 5's loop was. The difference is only what it is attached to — whatever Bazel invocation the command the user typed decides to make.

### `kinds` is not optional advice

```python theme={null}
ExecLogHook(on_entry = _spawn, kinds = ["spawn"])
```

One entry per spawn, plus one per file, directory, input set and runfiles tree, is 10k–50k entries on a medium build, and a Starlark call per entry is real money. `kinds` is applied in Rust before an entry ever becomes a Starlark value, and `["spawn"]` is commonly a few percent of the stream — 458 of this build's 8929 entries, 5.1%. Leave it off and you pay to convert all of it, then discard nineteen entries in twenty yourself. This is the filter step 5's task could not usefully have applied: its input-set reconstruction reads almost every kind there is. A spawn-shaped hook is the opposite case.

A spawn arrives with `id = 0` — every spawn does, because nothing refers to one. If you need to name or deduplicate a spawn, `@aspect//private/lib/execlog.axl` has a resolver that does it:

```python theme={null}
load("@aspect//private/lib/execlog.axl", "execlog")

res = execlog.resolver()
res.observe(entry)        # every entry, in on_entry
res.spawn_key(spawn)      # then, for a spawn: its primary output path
```

Resolution needs the entries that carry ids, so a hook that resolves must ask for `kinds = execlog.RESOLVER_KINDS`, not `["spawn"]` — which is also the point at which the `kinds` saving above stops being free.

## Your turn

Edit `config.axl` — not the command line — and set `--flaky_test_attempts=1`. Then run the test until the flake fires:

```shell theme={null}
aspect test //internal/domain/forecast:forecast_test --nocache_test_results
```

It fails about one run in three, and `--nocache_test_results` is what makes each run a fresh roll instead of a replay of the last cached pass. Eight attempts is usually enough to see it twice. When it goes, it is reported as a plain failure, your *flake* hook sees nothing, and the task exits 3:

```text theme={null}
  ⚙  EXECUTED: TestRunner×1

→ ❌ Failed test task (exit code 3) in 364ms · Tests failed
```

The `EXECUTED` line is still there because it comes from the other half of `config.axl`, the exec-log hook, which is watching actions rather than test verdicts. One spawn, because without retries Bazel runs the test once. It is the `FLAKE CAUGHT` line that is missing, and it is missing because without a retry there is no flake to catch — only a failure.

That is what the flag buys.

<Note>
  This one has to be changed in `config.axl`, and the reason is worth a minute. A trait's `extra_flags` are appended *after* the flags you typed, and Bazel takes the last value of a repeated flag — so your command line loses:

  ```shell theme={null}
  aspect test //internal/domain/forecast:forecast_test --flaky_test_attempts=1 \
    --announce-bazel-command=true 2>&1 | grep -o -- '--flaky_test_attempts=[0-9]*'
  ```

  ```text theme={null}
  --flaky_test_attempts=1
  --flaky_test_attempts=3
  ```

  Both are on the command line Bazel receives, and `=3` is the one that takes effect. Worth knowing before you spend an afternoon on a flag that "isn't working": in this repository, `config.axl` wins.
</Note>

Next: generate the targets nobody wants to hand-write.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.