With generate=True, RLRW samples a choice on every trial and learns from it, but it does not return that choice. Neither RLRW.export() nor Simulator.export() contains the simulated responses, so RLRW cannot be used to simulate data for parameter or model recovery.
The other built-in applications return their sampled choices: HybridMBMF returns action, PTSM returns chosen and PTSM2025 returns model_choice.
Reproduce
import numpy as np
import pandas as pd
from cpm.applications.reinforcement_learning import RLRW
from cpm.datasets import load_bandit_data
from cpm.generators import Simulator
data = load_bandit_data()
data["observed"] = data["response"]
np.random.seed(0)
model = RLRW(data=data[data.ppt == 1], dimensions=4,
parameters_settings=[[0.5, 0, 1], [5, 0, 10]], generate=True)
sim = Simulator(wrapper=model, data=data[data.ppt <= 3].groupby("ppt"),
parameters=pd.DataFrame({"alpha": [0.3, 0.5, 0.7], "temperature": [2.0, 5.0, 8.0]}))
sim.run()
print(list(sim.export().columns))
Output (cpm 0.26.0.dev0, with and without numba):
['policy_0', 'policy_1', 'reward', 'values_0', 'values_1', 'values_2', 'values_3',
'change_0', 'change_1', 'change_2', 'change_3', 'dependent', 'ppt']
There is no response column. The choices cannot be recovered reliably from the other outputs: reward is ambiguous when both arms give the same reward, and change is zero when the prediction error is zero.
Cause
cpm.applications._sessions.rlrw computes the choice on each trial (choice = choose(policy[t], uniforms[t]) if generate else response[t]), but returns only policy, reward, history, change, dependent. As a result, _RLRWModel.__call__ has no choice to put in its output.
Proposed fix
- Record the choice of every trial in
_sessions.rlrw, for example as an int64 array out_response, and return it.
- Add it to the output of
_RLRWModel as "response". That matches the column RLRW reads its observed choices from, so simulated data can be fitted directly.
- Test that the simulated responses are consistent with the returned policy, and identical with and without numba for the same
numpy.random.seed. This fits next to test_simulations_equal_the_per_trial_implementation and test_backends_agree in test/applications/test_backends.py.
Context
This came up in a benchmark of the hierarchical methods (100 simulated datasets × 100 participants), which fits RLRW with numba. Because of this issue, the choices had to be simulated with a separate implementation of the same model, and then checked against RLRW's policy.
With
generate=True,RLRWsamples a choice on every trial and learns from it, but it does not return that choice. NeitherRLRW.export()norSimulator.export()contains the simulated responses, soRLRWcannot be used to simulate data for parameter or model recovery.The other built-in applications return their sampled choices:
HybridMBMFreturnsaction,PTSMreturnschosenandPTSM2025returnsmodel_choice.Reproduce
Output (cpm 0.26.0.dev0, with and without numba):
There is no
responsecolumn. The choices cannot be recovered reliably from the other outputs:rewardis ambiguous when both arms give the same reward, andchangeis zero when the prediction error is zero.Cause
cpm.applications._sessions.rlrwcomputes the choice on each trial (choice = choose(policy[t], uniforms[t]) if generate else response[t]), but returns onlypolicy, reward, history, change, dependent. As a result,_RLRWModel.__call__has no choice to put in its output.Proposed fix
_sessions.rlrw, for example as anint64arrayout_response, and return it._RLRWModelas"response". That matches the columnRLRWreads its observed choices from, so simulated data can be fitted directly.numpy.random.seed. This fits next totest_simulations_equal_the_per_trial_implementationandtest_backends_agreeintest/applications/test_backends.py.Context
This came up in a benchmark of the hierarchical methods (100 simulated datasets × 100 participants), which fits
RLRWwith numba. Because of this issue, the choices had to be simulated with a separate implementation of the same model, and then checked againstRLRW's policy.