← Home
中文 Version

Why I Built DFTK Instead of Accumulating More Scripts and Tools

Read 5 min Words 1,691 Current Version English · 中文 Translation Available

Abstract

Traditional digital forensics tools are mature but rigid. They struggle to adapt to emerging AI-driven forensics scenarios involving large language models, agents, and local vector databases. Practitioners across the field commonly face the same pain points: disorganized tool collections, an overreliance on ad-hoc scripts, and a lack of standardized, reusable workflows.

This article introduces DFTK, an open-source modular digital forensics toolkit built for the AI era. It is designed to fill capability gaps left by traditional forensics tools and standardize the end-to-end workflow of AI-powered forensics analysis and evidence triage.

Keywords:AI · Agent · digital-forensics · DFTK · toolchain · tools

I have a folder on my disk named Toolkit. It’s crammed with all kinds of scripts and utilities. Ever since I got started with CTF and digital forensics, I’ve been dropping tools into it — from bundled tool packs bought on second-hand platforms to scripts and utilities I collected and downloaded on my own. The collection has only grown larger and more disorganized over time.

This folder looks nothing like the sleek, meticulously organized setups you see on seasoned tech professionals’ machines. It’s more like a storage closet that’s never been tidied: GitHub release downloads, quick-and-dirty Python scripts cobbled together with AI during competitions, .exe files shared by fellow forensics practitioners, a handful of .jar files whose purpose I long forgot, plus piles of folders named new, final and fix. At this point, I’d say half the contents are a total mystery to me.

What really bothers me isn’t how much disk space it takes up. It’s that same line of thought that pops up more and more often mid-competition:

I swear I had a script that does this.

What usually follows isn’t analysis — it’s a scavenger hunt for that script. Once I track it down, I then have to verify the Python version, install missing dependencies, read the README, and try to recall the command arguments. If I’m unlucky, the tool hasn’t been updated in years. If I’m really unlucky, it’s a script I wrote myself, back when I thought, “I don’t need to document this — I’ll definitely remember it in a few months.”

I never do, of course.

For a while, my “tool collection” looked something like this:

text
tools/
├── apk/
├── sqlite/
├── windows/
├── pcap/
├── misc/
├── misc2/
├── old/
└── ...../

The more tools I hoarded, the less reliable my actual capabilities became. What started as an attempt at organization gradually descended into chaos, until eventually I just started dumping everything in there.

Collecting code is not the same as building capability

Running a script successfully once is a world away from “having that skill at your disposal.”

The gap between “it worked once” and “I can reliably use this anytime” is not just one command — it’s an entire set of prerequisites you’ve long since forgotten.

Take SQLite, for example. I have half a dozen different tools for it: utilities for listing tables, inspecting WAL files, recovering deleted records, performing full-text search, plus a modified version I built for one specific challenge. But when faced with actual forensic evidence, I still have to figure out which tool to use, how to feed it input, and how to interpret its output.

Messy. Disorganized. Outdated. Those three words sum it up perfectly. EVTX parsers, Windows Registry scripts and everything else all suffer the same problem.

None of these tools are flawed on their own. The problem is they don’t interoperate — and a human has to sit in the middle acting as the glue:

text
File → Identify format → Find tool → Adjust parameters → Review results → Find the next tool

Doing this once is fine. Do it dozens of times, and the frustration builds up fast. That’s where the idea for DFTK was born. There was no grand manifesto about “building the next-generation digital forensics infrastructure.” I simply wanted to clean up the scripts and capabilities I actually use on a recurring basis.

My idea was straightforward. Start with the entry point. I can’t stand having to re-learn every six months: this tool uses -f, that one uses --input, and another puts the output directory as the second positional argument. The ideal isn’t that every feature uses identical command syntax — but at the very least, commonly used tools should have a consistent, discoverable interface, so you don’t have to do digital archaeology half a year later.

So I started building it.

A unified CLI is only the surface level

At first I thought standardizing the CLI would solve most of the problem. I was quickly proven wrong. The real challenge lies in the output.

The output from most forensics scripts is built for humans reading a terminal:

text
[+] Found URL: https://example.com/api

A human sees that line and immediately understands what it means, and can copy the URL to investigate further. A program cannot. At minimum, a program needs to know: which file the URL came from; whether it was extracted from a structured field, found via string scanning, or recovered through file carving; its exact offset within the file; whether the tool finished running; and if there’s no result, whether that means “there is genuinely nothing there” or “the parser failed to execute at all.”

Once you start asking those questions, bare print() statements simply don’t cut it.

I ended up spending a huge amount of time on DFTK not writing smarter parsers, but handling all the “unglamorous” fundamentals: result schemas, error states, data provenance, locators, and capability declarations. This kind of code never makes for impressive demos — there’s no flashy UI to wow people, and adding another error enum doesn’t make screenshots look cooler. But six months down the line, these are exactly the details that determine whether a tool can still be built upon.

For example: suppose a SQLite module returns zero records because an optional dependency wasn’t installed. If the upstream layer only receives:

json
{"success": true, "records": []}

That creates a serious problem. “No records found” and “the check never actually ran” are two completely different outcomes. A human watching the terminal would pick up on that nuance intuitively. An automated system will not.

Integrating with the Luduan Agent drove the point home

DFTK was never originally built for Agents. It wasn’t until I tried integrating DFTK with Luduan — the Luduan Digital Forensics Intelligent Agent — that I realized all those interfaces that “make perfect sense to a human” turn into severe problems when handed off to an Agent.

When a human sees a command return empty output, they might think: wrong scope, let’s try a different approach. When an Agent sees:

json
{"ok": true, "output": ""}

It will most likely only register the first part: success. Then it will proceed with the exact same strategy in the next step.

I ran into exactly this issue during a Luduan task. The Agent kept issuing actions, the tools all returned nominally successful, the terminal kept scrolling — it looked busy. But when I checked the final state files:

text
facts.jsonl       0
claims.jsonl      0
evidence.jsonl    0

In that moment, many of DFTK’s design choices suddenly had a much more concrete rationale. It’s not about being “AI-ready” — I don’t even particularly like that term. It’s simpler than that: if a tool can’t even explain “what just happened”, then anything built on top of it — whether human, script, or Agent — can only guess.

If a tool cannot clearly articulate what just happened, then everything above it — human, script, or Agent — is left to guesswork.

Not reinventing the wheel

Toolbox projects easily fall into a trap: if we’re standardizing everything, we might as well write everything from scratch. But I have no interest in that.

Tools like The Sleuth Kit, Volatility, and all those mature libraries for registry analysis, EVTX parsing, filesystem forensics and binary analysis already solve a host of hard problems. Rewriting them all just to have a “fully in-house” repository does little beyond introduce more bugs.

Instead, I’ve come to embrace writing glue code. It doesn’t sound as impressive as “framework”, but it’s accurate. If an existing capability works reliably, I adapt it in. If something is missing, I build it myself. What I maintain long-term is the outer layer of conventions: what inputs look like, how capabilities are discovered, how failures are described, how results can be referenced, and whether components can be composed together.

When a mature tool outputs a block of plain text, my job usually isn’t to rewrite its parser — it’s to wrap it in a robust, consistent adapter.

That also explains the key difference I care about now: the line between DFTK and just another scripts/ folder. If the project ended up as nothing more than:

text
dftk/
├── apk.py
├── sqlite.py
├── pcap.py
└── registry.py

Then all I’d have done is rename my old tools folder. What I want to add is everything that was missing between those scripts: stable identities, unified error handling, traceable results, capability discovery, and an interface callable from CLI, web interfaces, and Agents alike.

Not every module needs to be large. I actually prefer small, focused capabilities — each does one thing with clear boundaries, and the upper layer composes them as needed.

Ad-hoc scripts have become the new default

In the age of AI, many small tasks in AI decision workflows are written as temporary scripts. And honestly, humans should work the same way.

Sometimes grep is enough. Sometimes a dozen lines of Python is clearly faster than spinning up an entire toolchain. This doesn’t conflict with DFTK. I now separate my code much more deliberately: if it’s one-off for a single challenge, it gets discarded after use. If I run into the same problem a second or third time, and the input/output patterns have stabilized, then I consider folding it into the long-term toolset.

Not every piece of code deserves to be productized. And DFTK should not become just another, bigger code trash can.

That’s probably the clearest lesson I’ve taken away from this project. Before, when I found a good script, my instinct was to download it and save it. Now, if it actually solves a problem I keep running into, I ask myself one more question:

Will I still be able to use this capability directly, six months from now?

If the answer is still “hold on, let me go find the README first” — then it’s probably just another program I’ve collected, not a capability I actually possess. And whether that capability can be reused by an Agent is an even messier can of worms.

Closing

DFTK is now positioned as an Agent-native toolkit, built to fit the current wave of AI-enabled workflows.

I have prepared it for full integration with Agent frameworks.

If you’re interested, you can check it out here: DFTK on GitHub