The Behavioral Layer
BriefingsMapAboutRSS
the open note's neighborhood · Full map ↗

Tag: benchmark

9 items with this tag.

  • Aug 09, 2026

    OpenART Red-Teams Agents by Evolving the Environment Instead of the Prompt

    • red-teaming
    • evaluation
    • agent-harness
    • benchmark
    • stateful-environments
    • attack-surface
  • Aug 03, 2026

    A Validity Audit Finds Agent-Safety Scores Are Not Interchangeable

    • evaluation
    • benchmark
    • measurement-validity
    • agent-safety
    • capability
    • metrics
  • Jul 28, 2026

    Automatic Harness Evolution Fails to Beat Simple Test-Time Scaling

    • evaluation
    • agent-harness
    • overfitting
    • benchmark
    • scaffolding
    • test-time-scaling
  • Jul 20, 2026

    An IETF Draft Scores Agent Security on 55 Metrics

    • benchmark
    • standards
    • agent-security
    • excessive-agency
    • scoring
    • governance
  • Jul 09, 2026

    AgentHarm

    • agent-safety
    • misuse
    • jailbreaks
    • refusal
    • benchmark
  • Jul 09, 2026

    OSWorld

    • computer-use
    • agents
    • benchmark
    • evaluation
    • reliability
  • Jul 09, 2026

    WebArena

    • web-agents
    • benchmark
    • agents
    • evaluation
    • long-horizon
  • Jun 24, 2026

    ST-WebAgentBench

    • evaluation
    • safety
    • trust
    • policy-compliance
    • web-agents
    • benchmark
  • Jun 24, 2026

    tau-bench

    • evaluation
    • reliability
    • agents
    • rule-following
    • benchmark
    • behavioral-contracts

The Behavioral Layer · How this site is made · RSS

© 2026 Joel Goldfoot