Skip to content

RLRW(generate=True) does not return the choices it simulates #100

Description

@lenarddome

With generate=True, RLRW samples a choice on every trial and learns from it, but it does not return that choice. Neither RLRW.export() nor Simulator.export() contains the simulated responses, so RLRW cannot be used to simulate data for parameter or model recovery.

The other built-in applications return their sampled choices: HybridMBMF returns action, PTSM returns chosen and PTSM2025 returns model_choice.

Reproduce

import numpy as np
import pandas as pd
from cpm.applications.reinforcement_learning import RLRW
from cpm.datasets import load_bandit_data
from cpm.generators import Simulator

data = load_bandit_data()
data["observed"] = data["response"]
np.random.seed(0)

model = RLRW(data=data[data.ppt == 1], dimensions=4,
             parameters_settings=[[0.5, 0, 1], [5, 0, 10]], generate=True)
sim = Simulator(wrapper=model, data=data[data.ppt <= 3].groupby("ppt"),
                parameters=pd.DataFrame({"alpha": [0.3, 0.5, 0.7], "temperature": [2.0, 5.0, 8.0]}))
sim.run()
print(list(sim.export().columns))

Output (cpm 0.26.0.dev0, with and without numba):

['policy_0', 'policy_1', 'reward', 'values_0', 'values_1', 'values_2', 'values_3',
 'change_0', 'change_1', 'change_2', 'change_3', 'dependent', 'ppt']

There is no response column. The choices cannot be recovered reliably from the other outputs: reward is ambiguous when both arms give the same reward, and change is zero when the prediction error is zero.

Cause

cpm.applications._sessions.rlrw computes the choice on each trial (choice = choose(policy[t], uniforms[t]) if generate else response[t]), but returns only policy, reward, history, change, dependent. As a result, _RLRWModel.__call__ has no choice to put in its output.

Proposed fix

  • Record the choice of every trial in _sessions.rlrw, for example as an int64 array out_response, and return it.
  • Add it to the output of _RLRWModel as "response". That matches the column RLRW reads its observed choices from, so simulated data can be fitted directly.
  • Test that the simulated responses are consistent with the returned policy, and identical with and without numba for the same numpy.random.seed. This fits next to test_simulations_equal_the_per_trial_implementation and test_backends_agree in test/applications/test_backends.py.

Context

This came up in a benchmark of the hierarchical methods (100 simulated datasets × 100 participants), which fits RLRW with numba. Because of this issue, the choices had to be simulated with a separate implementation of the same model, and then checked against RLRW's policy.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions