As organizations increasingly adopt generative artificial intelligence tools for software development, incident responders face new challenges when investigating environments where AI agents operate. To address this visibility gap, two new Python scripts have been released to assist security professionals in reconstructing AI agent activity during forensic engagements, as detailed by the Internet Storm Center.
Investigating AI Coding Assistants
Modern development environments frequently incorporate AI coding assistants and autonomous agents such as Claude Code, Codex, Gemini CLI, Cursor, Copilot, Warp, Windsurf, Qwen Code, OpenCode, and Hermes. These tools generate extensive logs, chat histories, model usage metrics, and API request payloads that can provide critical context during a security assessment or breach investigation.
To simplify the extraction and analysis of this evidence, new forensic tooling has been developed. The opencode-chat-replay.py and hermes_forensic_extract.py utilities target artifacts left behind by OpenCode and Hermes respectively, allowing analysts to examine what prompts were submitted, what responses were returned, and what underlying tool calls were executed.
Technical Details of OpenCode Replay
The opencode-chat-replay.py script parses SQLite databases associated with OpenCode sessions. It supports both legacy separate message tables and newer consolidated session message storage. Key technical characteristics include:
- Snapshot Integrity: Prior to querying, the script copies the SQLite database along with its
-waland-shmsidecars into a temporary directory, using SQLite read-only URI mode to prevent any modification of original forensic evidence. - Flexible Output Formats: Generates Markdown transcripts with collapsible detail blocks for reasoning and tool calls, standard JSON object exports, and JSONL formats for line-by-line processing.
- Granular Filtering: Supports
--startand--endtime parameters, session listing, slug matching, and child-session inclusion.
Hermes Forensic Extraction
The hermes_forensic_extract.py script takes a broader extraction approach across multiple artifact locations, targeting:
~/.hermes/state.dbfor session tracking, messages, and model usage metrics.~/.hermes/sessions/request_dump_*.jsonfor full LLM API request and response payloads.~/.hermes/logs/*.logfor gateway, GUI, and agent error logs.
Output from the Hermes extractor is structured as newline-delimited JSON, giving investigators a comprehensive view of surrounding context and API interactions.
Recommendations for Defenders
When conducting incident response on developer workstations or cloud-hosted build servers where AI agents are active, security teams should incorporate these artifacts into their standard evidence collection scope.
- Preserve Sidecar Files: Ensure disk acquisitions capture SQLite WAL and SHM files alongside primary database stores to prevent data loss.
- Audit for Secrets: Remember that while authentication tokens may remain untouched, recorded tool inputs and outputs within AI session logs frequently contain sensitive credentials or internal source code.
- Use Read-Only Procedures: Always execute forensic extraction tools against copied evidence images rather than live production environments to preserve chain of custody.
Related: AI-Driven Vulnerability Surges and UAT-11795 Starland RAT Campaign, Emerging Cyber Threats and Espionage Risks in Neurotechnology