Bughunter · Security Research · Digital Forensics & OSINT
Offensive Security · AI systems · agent tooling
I break systems I actually use, trace failures to their root cause, and turn the evidence into fixes, tools, or upstream reports.
My current focus is offensive security around AI and developer tooling: real systems under real load, reproducible failures, preserved artifacts, and attribution that survives a second look.
Public lab: AI metrology · digital forensics · agent memory · security research · FS25 engine archaeology
- BRONCO — FC-001 · policy-driving evaluation provenance & reproducibility.
- schroedinger-sync — local-first AI data sovereignty · hardening under real workloads.
- MemPalace — upstream fixes and persistent-memory work from production use.
19 Aug— commented on issue EleutherAI/lm-evaluation-harness#374919 Aug— commented on issue SWE-bench/SWE-bench#62119 Aug— commented on issue KeilerHirsch-Labs/BRONCO-AI-Metrology-Benchmarks-DIN-ISO-IEC#1
Auto-refreshed from public GitHub activity every 6 hours. Profile-maintenance commits are excluded.
- Courseplay/Courseplay_FS25#1298 — fix: don't build the vehicle HUD on a dedicated server
- modelscope/evalscope#1587 — feat(metrics): Wilson CI + paired McNemar helpers
- tobi/qmd#901 — fix(embed): report actual completion and session expiry (additive)
- MemPalace/mempalace#1989 — fix(embedding): preload CUDA/cuDNN DLLs before the first GPU session
Selected automatically from my currently open public PRs, preferring one active contribution per upstream repository.
- KeilerHirsch-Labs/schroedinger-sync
v2.2.0— published 19 Jul 2026 · SHA256SUMS
- Threat model / security policy — constraints, caveats, reporting path, and what the tests actually enforce.
- BRONCO Field Cases — evidence-linked measurement cases with falsifiers and explicit open questions.
- Release checksums are linked beside shipped releases when available.
- schroedinger-sync — local-first export and sync for
claude.aiconversations, project docs, and memory. Windows, DPAPI + CDP, no telemetry, no cloud. - BRONCO — AI Metrology, Benchmarks, DIN & ISO/IEC — research-first infrastructure for treating AI benchmarks as versioned measurement instruments, models/agents/harnesses as precisely defined systems under test, and results as measurements with provenance, applicability, and uncertainty.
- AFFLR_Anthropic-Failure-Forensics-Live-Radar — watches the public
anthropics/claude-codeissue space and prioritizes security/trust-boundary, evidence/provenance/integrity, and fresh critical signals for human review. - MemPalace/mempalace — my daily persistent-memory layer for agent workflows. I contribute fixes found by running it under real-world load.
- Authorized targets. Offensive mindset.
- Hypothesis → reproduction → evidence → root cause → fix.
- Measure, don't guess.
- Evidence before attribution.
KeilerHirsch-Labs · MyRank.dev · Ko-fi
📫 Open an issue on the relevant repository — that's the reliable way to reach me.

