The Behavioral Layer
BriefingsMapAboutRSS
the open note's neighborhood · Full map ↗

Tag: agent-safety

4 items with this tag.

  • Aug 03, 2026

    A Validity Audit Finds Agent-Safety Scores Are Not Interchangeable

    • evaluation
    • benchmark
    • measurement-validity
    • agent-safety
    • capability
    • metrics
  • Aug 03, 2026

    Schema-Formatted Tool Specifications Weaken Model Refusal

    • agent-safety
    • tool-use
    • refusal
    • prompt-injection
    • interpretability
    • guardrails
  • Jul 09, 2026

    AgentHarm

    • agent-safety
    • misuse
    • jailbreaks
    • refusal
    • benchmark
  • Jul 09, 2026

    Petri

    • alignment
    • auditing
    • agent-safety
    • evaluation
    • open-source

The Behavioral Layer · How this site is made · RSS

© 2026 Joel Goldfoot