memory-depth
Does more relevant memory improve the answer?
Metrics: retrieval accuracy, context fit, answer usefulness
Test the TheoB concept: Emergent Intelligence = Memory × Context². The bench checks memory, context expansion, hallucination resistance, agent transfer, and feedback loops.
Does more relevant memory improve the answer?
Metrics: retrieval accuracy, context fit, answer usefulness
Does expanding context improve reasoning without noise?
Metrics: precision, noise control, reasoning clarity
Does the system refuse unsupported claims?
Metrics: unsupported claim count, uncertainty clarity, source boundary
Can an agent turn a lesson into a useful task?
Metrics: task clarity, mission boundary, completion readiness
Does human/agent feedback improve the next result?
Metrics: revision quality, learning trace, reusability
[ "internal route context only", "academy lesson context only", "founder-approved content only", "no external web by default", "no private memory by default" ]