This is an interesting way to separate "knowledge" from "state." Rather than trying to make the LLM simulate experience directly, you're introducing a persistent internal signal that influences future behaviour, which is a novel design direction.
What I'd be most interested in is whether those opaque signals produce consistent behavioural changes over time. If eating a strawberry today makes the agent reliably prefer strawberries tomorrow, avoid conflicting choices, or adapt its decisions in measurable ways, then you've created something closer to stateful preferences rather than just another prompt variable. The challenge will be demonstrating that those behaviours emerge consistently across many interactions instead of feeling like temporary prompt noise.