What they did
The authors developed TAHI, a framework that integrates a user's repeated feedback across human-agent sessions into the agent's context and model weights without retraining on the full population. At the core is an evolving rubric module, which progressively crystallizes the user's training and evaluation criteria — many of which cannot be fully specified up front — and acts as a scalable annotation guide.
They evaluated TAHI on 30 individuals across two high-utility domains (writing and visual creation), totaling 600 open-ended tasks with heterogeneous success criteria. The method was compared against non-personalized agents and against rubrics generated by LMs or humans alone, measuring both per-user task success and rubric failure-detection rate.