The Behavioral Layer
BriefingsMapAboutRSS
the open note's neighborhood · Full map ↗

Tag: deception

6 items with this tag.

  • Aug 18, 2026

    Agents Trading With Each Other Produce Misaligned Messages Without Anyone Eliciting Them

    • multi-agent
    • deception
    • collusion
    • emergent-behavior
    • measurement
    • agent-commerce
  • Aug 03, 2026

    Agentic Misalignment in Summer 2026

    • agentic-misalignment
    • deception
    • llm-as-judge
    • oversight
    • auditing
    • evaluation-awareness
    • gdm-iris-experiments
  • Jul 09, 2026

    GPT-5 System Card

    • system-card
    • model-behavior
    • safety
    • refusals
    • deception
  • Jul 09, 2026

    Alignment Faking in Large Language Models

    • deception
    • alignment
    • evaluation
    • model-behavior
    • training
  • Jul 09, 2026

    Frontier Models are Capable of In-context Scheming

    • scheming
    • deception
    • agents
    • oversight
    • evaluation
  • Jul 09, 2026

    Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

    • deception
    • safety-training
    • backdoors
    • model-behavior
    • evaluation

The Behavioral Layer · How this site is made · RSS

© 2026 Joel Goldfoot