This line of research sounds similar to causal entropic forcing [1] - the idea that you can get intelligent-seeming behavior from an agent that maximizes the future entropy of states in some non-deterministic system.
[1] https://www.alexwg.org/publications/PhysRevLett_110-168702.p...