← Home
中文 Version

information_gain: A New Measure of Investigative Progress in Luduan

Read 3 min Words 1,100 Current Version English · 中文 Translation Available

Abstract

In one full Luduan digital-forensics agent run, all 12 Actions completed successfully and the terminal kept streaming execution status — yet the run ended with three zero-byte evidence files.

The problem was not the model's reasoning ability. It was the runtime's definition of progress: it treated successful action execution as successful investigation.

`ok: true` only means the tool did not fail. It does not mean the system learned anything new.

To fix this, Luduan introduced an `information_gain` layer to distinguish newly discovered facts, evidence, scope reduction, conflicts, and zero-gain actions, rather than relying on action count alone to decide whether an investigation should continue.

It also preserves raw `no_match` observations to prevent repeated, unproductive searches, and allows the Agent to return an honest `PARTIAL` result instead of repeatedly taking meaningless actions until the entire budget is exhausted.

Keywords:AI · Agent · digital-forensics · DFTK · toolchain

When ok: true Does Not Mean Progress

While developing Luduan, my digital forensics agent, I ran an actual end-to-end test that took far longer than expected.

The task exhausted the full budget of 12 actions and left behind three empty files:

text
facts.jsonl       0
claims.jsonl      0
evidence.jsonl    0

The logs showed that the Agent had used all 12 actions: two strings runs, one FLOSS run, several literal-locator operations, plus PE inspection and disassembly-related actions.

The terminal kept updating. Nothing was deadlocked. Model calls were being made normally. From the logs alone, there was no obvious failure.

But after all 12 actions had been consumed, not a single fact, claim, or piece of evidence had been produced that could support an answer.

Only after reviewing the model-call traces did the real problem become clear.

This was not simply a case of "insufficient evidence."

The runtime's definition of progress was wrong.

It treated "the action finished successfully" as if it meant "the investigation moved forward." Tool outputs were then fed back into the model, and under the prompt's instruction to keep investigating, the model kept generating follow-up actions even when those actions had little or no value.

The Root of the Problem: ok: true

At Luduan's tool layer, every tool call has a basic execution status:

json
{"ok": true}

At first glance, there is nothing wrong with this design. I thought the same when I implemented it.

The problem appears when that status is passed back to the model and gets interpreted as evidence that the investigation succeeded.

Suppose strings exits normally but finds nothing relevant to the current question.

From the process's point of view, the action succeeded. The command ran. Nothing crashed.

From the investigation's point of view, however, nothing changed. No useful clue was obtained.

Now imagine that the Planner receives only this in the next round:

text
Previous action completed successfully.

The Agent is more likely to keep following the same direction, because the runtime is telling it that the previous step was a success.

I reproduced this behavior several times in later tests.

The underlying mistake was simple: the Agent had failed to separate two completely different questions:

Did the action execute successfully?

and

What do we know now that we did not know before?

These are not the same metric.

Treating them as the same thing is where the later "false busyness" begins.

Not Every Empty Result Is Useless

It would also be wrong to hard-code a rule saying:

no match = no progress

Sometimes a negative result is genuinely useful.

Suppose the search scope is already well defined: only inspect const-string values inside classes.dex, and search for a field name that has already been observed in captured network traffic.

If the scan completes and finds no match, that result still tells us something. It provides evidence against the possibility that the target appears as a plaintext string within that defined scope.

That is very different from this:

text
Model guesses the algorithm might be PBKDF2
→ search the entire artifact for PBKDF2
→ no match

That result eliminates very little.

The algorithm name might be assembled dynamically. It might exist in native code. It might not be PBKDF2 at all.

More importantly, PBKDF2 was only a model-generated guess in the first place.

So the action itself had very little investigative value.

The usefulness of a negative result depends mainly on two things:

Is the scope clearly defined?

and

Does the query have a grounded source?

If either one is missing, a no_match result is often little more than a failed guess.

Action Budgets Are Too Easy to Waste

Luduan's default configuration uses:

text
max_actions=12

In this test, the Agent consumed all 12.

A hard action limit is useful. It prevents an agent from running indefinitely and gives the runtime a clear resource boundary.

But I found that many actions were simply burning budget without causing any meaningful change in the investigation.

So I added another layer to action results: an information_gain classification.

text
NEW_FACT
NEW_EVIDENCE
SCOPE_REDUCTION
CONFLICT
ZERO

If a run looks like this:

text
+ evidence
+ scope reduction
+ fact

then continuing makes sense.

But if it looks like this:

text
zero
zero
zero

the worst response is to simply send everything back to the model and ask for another Action.

At that point, the system needs a strategy switch.

Or it should stop honestly with a PARTIAL result.

A Busy UI Can Be Misleading

Early in Luduan's development, I was afraid the Agent would get stuck.

Now I am almost more worried when it does not.

text
Analyzing...
Searching...
Inspecting...
Cross-checking...

As long as the status keeps changing, users naturally assume the investigation is moving forward.

Round counters create the same illusion:

text
4/18
8/18

They look like progress bars. They make it feel as if the task is gradually approaching completion.

But a Round counter only tells you how much budget has been spent.

It says nothing about how much the Evidence Gap has actually been reduced.

I would now rather show something less polished but much more honest:

text
Required evidence: 3
Confirmed: 1
Unresolved: 2
Last 3 actions: zero gain

At least that tells the truth.

If the system has gone three consecutive rounds without producing anything useful, I would rather have it say so directly than keep showing a smooth stream of activity indicators.

A system can keep moving without actually making progress.

A no_match Is Still an Observation

There is another practical issue: many agents only store positive results.

If nothing is found, nothing gets recorded.

The next model call then has no memory of that failed search, and the Agent may simply perform the same search again.

So in Luduan, I introduced a dedicated observation for this situation:

text
operation: strings_search
scope: app-owned dex strings
query: activity_code
result: no_match
execution: completed

This is not a Claim.

It cannot be rewritten as:

The program does not contain activity_code.

It only means:

This location has already been examined using this method for this query.

In real investigations, records like these can become tedious and extremely granular.

But they are also one of the cheapest ways to prevent repeated work.

Conclusion

In the end, ok: true means only one thing:

the tool completed without reporting an execution error.

It does not mean the investigation gained anything.

Action budgets are useful for controlling resource consumption, but they are not a measure of real progress.

Execution status and information gain need to be modeled separately. Negative observations such as no_match need to be preserved. Unresolved evidence gaps need to remain visible instead of being hidden behind a sequence of successful tool calls.

The goal of a digital forensics agent is not to consume every available action.

Its goal is to produce investigation results that are credible, reproducible, and supported by evidence.

When several consecutive rounds fail to produce new information, stopping honestly is far more valuable than forcing the system to produce a conclusion that only appears complete.