Credentials
Press commentary, peer review, research data and standards work.
Cited in the press
Expert commentary in the technology press.
Going Offline Doesn't Remove Password Risk. It Swaps It.
Expert commentary: offline versus cloud password managers — which failure mode you choose
Anthropic Discloses Fourth Unauthorised Claude Access Incident: Is The Security Industry Prepared For AI Breaches?
Expert commentary: AI agents with legitimate access and enterprise threat models
Peer review
Open peer review for Qeios and PREreview, each review with its own DOI.
- 2026-09-14·PREreview·Decision-making under uncertainty
- 2026-09-13·PREreview·Evaluation of medical language models
The widening evaluation gap in medical large language model research 2023 to 2026
- 2026-09-12·PREreview·Reproducibility of agentic systems
- 2026-09-12·Qeios·Domain-specific financial language model
- 2026-09-12·Qeios·Identity drift in LLM agents
Prompt Volatility: An Empirical Study of Identity Drift in LLM Agents
- 2026-09-12·Qeios·Dynamic routing between models
- 2026-09-12·Qeios·Machine-assisted mathematical proofs
- 2026-09-12·Qeios·Cognitive bias in model outputs
- 2026-09-12·Qeios·Energy management in the home
- 2026-09-12·Qeios·Cultural alignment of models
- 2026-09-12·Qeios·Long-form opinion summarisation
Research and open data
Technical reports and the datasets behind them.
Clear Every Third Task: A Measured U-Curve in the Context Economy of Coding Agents
How often should an agent's context be cleared? 36 runs, six session-length policies, six replicates each, 4086 tests and zero failures. Cost is U-shaped in session length: clearing after every task costs 25.3% more than clearing every third, and never clearing costs 15.6% more. Exact permutation tests put both extremes outside the optimum (p = 0.0022 and p = 0.0108) while three, four and six tasks per session are indistinguishable - the optimum is a plateau, not a point. The U decomposes into cache writes falling and context-per-call rising, priced 12.5:1 against each other.
The Context Economy of Agentic LLM Sessions: Where the Money Actually Goes
722 agent sessions, 150,902 model calls, 34.6 billion tokens. Context handling accounts for 83.5% of modeled cost and generation for 16.5%; 80% of the spend comes from 3.3% of sessions. Includes two log-deduplication traps that change the answer by a factor of two in either direction, and the de-identified dataset.
AI Tools Radar: GitHub and Hugging Face projects selected by arsentev.ai (monthly)
490 open-source AI projects selected and reviewed by the arsentev.ai radar since June 2026, with daily GitHub and Hugging Face API snapshots of how they evolve after selection. A new version is released every month.
Context U-curve: 36 coding-agent runs under six context-clearing policies
Run-level token counters, modeled cost, wall clock and test outcomes for every run behind the U-curve report. Counters only - no prompts, no model output, no paths.
Standards
Internet-Drafts at the IETF.
Agent Run Metrics: A JSON Interchange Format for Resource Accounting of AI Agent Runs
draft-arsentev-agent-run-metrics
Discovery and Retrieval of Publisher-Curated Context Files for Large Language Models
draft-arsentev-llm-context-discovery
Open-source tools
contextburn — a context-accounting meter for coding agents, MIT licensed.
Group participation
Standards working groups on AI agents.
Research — peer-reviewed papers and the author's identifiers. · Press · About