
“Just-in-Time” adaptive interventions (JITAIs) can provide daily step goals and financial incentives (e.g., reward points) to enable participants to reach healthy levels of physical activity. However, personalization and optimization remain challenging because participant responses vary over time and between individuals. Recent JITAI efforts have benefited from control-oriented strategies using system identification and Model Predictive Control (MPC), with YourMove (NCT05598996) providing an illustrative example. MPC offers a structured decision framework for selecting feasible goal trajectories from theory-guided dynamic behavioral models, but its closed-loop performance depends on tuning choices that trade off tracking speed, smoothness, constraint satisfaction, and other criteria, making manual selection difficult. A contextual multi-armed bandit formulation can instead use a reward function to select MPC tuning parameters that best align the controller with the intervention outcomes of interest, conditioned on observed participant context. This paper presents a simulation-based proof of concept in which a TD3-based agent is used as an outer-loop tuner that learns the optimal MPC tuning parameters, while MPC produces feasible responses that satisfy intervention constraints. Results demonstrate improved closed-loop behavior tracking and increased self-efficacy relative to a fixed-tuning MPC baseline, and multi-seed training indicates reproducible convergence to high-reward tuning policies. Both novel and practical insights on the synergism of machine learning and control engineering are outcomes of this research.