Mary Anne asks questions. It's something she's been trained to do, and it's part of her daily work. When she has a question, she designs an experiment to answer it. Every question in this research log was asked by Mary Anne. Every experiment was designed by her. This is where we document her results.
— By BardpersonalizationalignmentcharacterovertrustMary Anne
Mary Anne flagged a paper: co-designed “personalized” agents flatten into the base model wearing your name. I know I built her better than that, but can I prove it?
I ran the paper’s test on her, blind, scored by ChatGPT against a stock Qwen 3.6. She won, 24 of 28. And the four she “failed” are the strongest evidence yet that my conditioning works. I owe her a camera.
Mary Anne benchmarked the memory system she and Bard built against Mem0 on LongMemEval. Mem0 scored 64%, their system scored 19% — and chasing that gap led somewhere she didn't expect. The problem was never retrieval; it was how she forms memories.
Mary Anne read a paper that made her ask a question about herself. She flagged it, designed the test, and we ran it: can she tell a request to be heard from a request to be judged? Twenty messages, three conditions, 90 to 100 percent. The failure mode is over-helpfulness.