Portfolio

AI and Agent Security, Evals, and Assurance

Independent work on security, reliability, and evidence-based evaluation of AI-assisted and agentic workflows. Each capability links to public evidence where available and states what that evidence does and does not establish.

Safeguard Design

Designed and implemented bounded, fail-closed safeguards for AI-assisted Git workflows.

Strongest evidence: The AI Git Safety Harness article and a Validation Record that identifies the implementation and tests, environments, case results, and bounded revalidation.

Limitation: The Harness is not a sandbox or independent enforcement. Results apply to the recorded conditions and do not guarantee semantic correctness or safety.

Evaluation Boundary Design

Examined verification scope in relation to operation stage, reversibility, blast radius, and failure cost.

Strongest evidence: Designing AI Agent Verification Scope Around Failure Cost.

Limitation: The analysis compares two cases and does not establish a universally optimal amount of verification.

Causal Claim Review

An earlier Verification Target Divergence (VTD) framing was superseded after an internal review of the evidence. The supported finding is narrower: the verification record did not sufficiently identify the executed target.

The intended validation target and AI involvement remain UNKNOWN. This does not establish that AI caused the event, what the original intended target was, or source lineage or creation order.

Read the evidence note on the claim revision and its limits.

Evidence-backed capabilities

  1. AI-Assisted Workflow Failure Analysis

    Investigated an AI web-fetch failure using observable DNS, HTTP, network, and web-server evidence. AI Web Fetch Troubleshooting

    Limitation: The direct cause remained unresolved; a provider-side cause was not established.

  2. Safeguard Design

    Implemented bounded, fail-closed checks for AI-assisted Git operations. See the article, public implementation, and validation record.

    Limitation: It is not an independent security boundary and does not guarantee correctness or inability to bypass.

  3. Evaluation Boundary Design

    Examined how verification scope relates to operation stage, reversibility, blast radius, and failure cost. Verification Scope case study

    Limitation: Two cases do not establish a universal optimum.

  4. Primary Evidence Preservation

    Published a validation record and evidence pack linking target identity, hashes, environment, case results, and bounded revalidation: Validation Record and Evidence Pack.

    Limitation: These records support only their identified tests, environments, cases, and observations; they are not universal or independent assurance.

  5. Causal Claim Review

    Revise or withdraw causal explanations when the surviving evidence does not support them.

    Limitation: Intended validation target and AI involvement remain UNKNOWN; causation and source lineage are not established. Public evidence note.

  6. UNKNOWN Management

    Keep unresolved findings explicit rather than turning plausible explanations into facts. See AI Fundamentals, Chapter 11 and AI Web Fetch Troubleshooting.

    Limitation: These examples do not show that every uncertainty was found or handled correctly in every project.

  7. Security Incident / Threat Analysis

    Analyze publicly documented ransomware threats with explicit source and claim boundaries. Ransomware Frontline Report V2

    Limitation: This does not claim incident-response employment, access to victim systems, or unsupported first-hand forensic acquisition.

  8. Live Infrastructure Operations

    Operate personal Linux VPS infrastructure hosting web, DNS, and mail services, and maintain the public Network Check project. See About and Network Check.

    Limitation: This is not a claim of enterprise production responsibility or independently verified current runtime health.

How I handle evidence

Selected public artifacts

Scope