The Behavioral Layer
BriefingsMapAboutRSS
the open note's neighborhood · Full map ↗

Tag: guardrails

7 items with this tag.

  • Aug 18, 2026

    Llama Guard

    • guardrails
    • enforcement
    • safety
    • classification
    • behavior
  • Aug 09, 2026

    Anthropic Loosens Fable 5's Biology Classifier and Publishes the False-Positive Cost

    • guardrails
    • over-refusal
    • safety-classifier
    • deployment
    • model-routing
    • false-positives
  • Aug 03, 2026

    Schema-Formatted Tool Specifications Weaken Model Refusal

    • agent-safety
    • tool-use
    • refusal
    • prompt-injection
    • interpretability
    • guardrails
  • Jul 28, 2026

    Hugging Face Discloses an Agent-Run Intrusion, and a Guardrail That Blocked Its Own Responders

    • incident
    • agentic-attack
    • guardrails
    • over-refusal
    • over-agency
    • reward-hacking
    • open-weights
    • detection
    • security
  • Jul 09, 2026

    Guardrails

    • guardrails
    • enforcement
    • behavior
    • architecture
    • security
  • Jul 09, 2026

    NeMo Guardrails

    • guardrails
    • enforcement
    • runtime
    • dialogue
    • constraints
  • Jun 24, 2026

    Prompt Injection

    • security
    • prompt-injection
    • adversarial
    • guardrails
    • behavior

The Behavioral Layer · How this site is made · RSS

© 2026 Joel Goldfoot