NeurIPS 2026 · Proposed Workshop

Hidden Information in AI Outputs

What info hides in model outputs? How does it get there? How do we tell attack from defense? Provenance meets oversight.

About

Models may hide information in their reasoning traces and outputs in ways that evade monitors and oversight pipelines. Hidden information also includes what models inadvertently disclose: memorized training data recoverable from outputs and proprietary model internals extractable through query access.

Watermarks embed a recoverable signal for provenance, while steganographic collusion and monitor evasion embed signals to escape oversight. These are formally the same object viewed from opposite sides — a defense for one is an attack on the other — yet the communities working on them rarely speak to one another.

This workshop brings the provenance side (watermarking) and the oversight side (steganography detection, monitoring, control, and interpretability) together, alongside the cryptography, privacy, and infrastructure communities that share the problem.

Topics of Interest

  1. Taxonomization. Proactively mapping the space of hidden reasoning. Existing maps cover encoded reasoning and multi-agent collusion channels, but none yet spans the provenance side.
  2. Defense, detection, and mitigation. Detecting whether generated text contains a covert signal, distorting output to strip such signals, and conversely making benign hidden payloads resilient to adversarial removal.
  3. Mechanistic interpretability. Structural interventions that incentivize or ablate information-hiding behavior. Probes and SAEs that localize encoding and decoding features in later-layer activations can already recover hidden secrets and objectives.
  4. Emergence. Whether covert encoding strategies emerge under RL optimization pressure. Sub-tracks include reward-hacking case studies; multi-agent collaboration, supervision-aware and unaware; testing emergent encodings against catalogued ones; and adjacent covert behaviors such as sandbagging under monitoring.
  5. Theory. Information-theoretic guarantees that can guide future monitors and defense strategies. The steganographic-security literature is now being connected directly to LLM monitoring.
  6. Infrastructure. Integrating AI control, watermarking, and monitors into real-world deployments. Open topics include trusted execution environments, the monitoring "tax", auditability of reasoning traces, and open tooling for control and watermark verification.

More information forthcoming. The full program — invited speakers, panel, schedule, organizers, and call for papers — will be posted here as it's finalized. For inquiries in the meantime, contact matthew@poseidonresearch.com.