Changelog
All notable changes to POMDPPlanners are documented here, newest release first. This project adheres to Semantic Versioning.
Each #NNN links to the pull request that introduced the change.
Release 0.5.0 (WIP)
CARLA, nuPlan and Isaac Lab environments; vectorized planning and batched GPU beliefs
Breaking Changes:
Environments now own their visualization file naming. Callers that built visualization paths themselves must go through the environment (#203).
New Features:
Added
CarlaPOMDP, a forward-only world environment backed by the CARLA simulator (#204).Added a pluggable generative-model interface and a multi-agent perception world for CARLA (#207).
Added planner-side perception in the CARLA world observation model, so the planner model and the world emit identical raw observations while the belief filters and infers hidden per-agent intent (#208).
Added a destination/route option for CARLA with a success terminal and route metrics (#221).
Added a headless CARLA server pool for parallel episode simulation (#217).
Added a nuPlan POMDP environment — world, model, perception and belief — with a POMCPOW example (#210).
Added an Isaac Lab POMDP environment and visualizer (#206).
Added a vectorized belief tree to
POMDPPlanners.core(#214).Added the VOPP vectorized planner and vectorized generative models (#215, #216).
Added
BatchedParticleBelief, a torch-backed belief for batched GPU belief filtering (#218).Added opt-in
is_*_hit_terminalhazard-terminates-episode flags across the hazard environments, with a draw-coupled termination gate (#211, #212).Added
EVADE_WHEN_SPOTTEDandPURSUEopponent-policy support to the vectorized LaserTag model (#219).
Bug Fixes:
Documentation:
Others:
Bumped the package version to 0.5.0 across
pyproject.toml,POMDPPlanners/__init__.py,test_setup.pyand the Sphinxreleasestring, which had been stale at 0.2.0 (#228).Refactored CARLA onto per-channel observation models with an
encode_observationseam (#209).Added coverage for the LaserTag belief updater honouring
opponent_policy(#202) and for CARLA observation equality and collision-penalty reward (#205).
Release 0.4.0 (2026-05-28)
Constrained planners, hazard-centric reward variants, run-progress notifications
Breaking Changes:
New Features:
Added the constrained planners CPOMCPOW and CPFT-DPW together with a
ConstrainedEnvironmentABC (#170).Added run-progress tracking: Slack notifications on run start, finish and failure, backed by a SQLite progress DB with an external stall-detection watcher for deaths the process cannot witness itself (#172).
Added per-task progress callbacks for the Dask and PBS backends (#173).
Added opt-in circular dangerous areas to
PacManPOMDP(#176), plus a dangerous-area step counter and a total-danger-encounters aggregate (#187).Added high-variance and decaying reward variants to RockSample — with batch acceleration — Push and PacMan (#180, #181, #182).
Added
to_unique_support_distributiontoVectorizedWeightedParticleBelief(#188).Added a selectable LaserTag
OpponentPolicywith evade, pursue and evade-when-spotted behaviours (#196, #197).
Bug Fixes:
Fixed the LaserTag opponent to evade and react to the robot’s pre-move position (#195).
Aligned the C++ reward kernels with their Python counterparts across five environments (#183) and closed Python-side reward-kernel bugs in PacMan and RockSample (#184).
Fixed three environment correctness bugs in the sanity, Tiger and LightDark environments (#185).
Fixed the PacMan visualizer to render the ghost belief heatmap for vectorized particle beliefs (#174).
Documentation:
Others:
Replaced the deprecated
pkg_resourceswithimportlib.metadata(#171).Extracted generic dangerous-area Numba kernels (#177) and reward models for RockSample, Push, LaserTag and PacMan (#178, #179).
Added MIT SPDX headers to all source files (#189).
The release workflow now builds an sdist only; the previous wheel carried an unportable
linux_x86_64platform tag that PyPI rejects (#193).Made environment tests independent of constructor defaults (#199).
Release 0.3.1 (2026-05-12)
Packaging hotfix for a broken 0.3.0 sdist
Bug Fixes:
Bundled the C++ headers in the sdist. Without them
pip installfrom PyPI failed to compile, becausesetup.py’sinclude_dirspointed at paths absent from the tarball. Added packaging regression tests (#168).
Others:
Release 0.3.0 (2026-05-11)
Risk-sensitive benchmarking mechanics, arena tree, large performance and correctness pass
Breaking Changes:
Environment.rewardgained an optionalnext_stateparameter so penalty terms (obstacles, dangerous areas) are scored against the realised post-transition state rather than a fresh sample. Threaded through rollout and POMCP-DPW.PacMan state is now a raw NumPy
ndarray; the native batch path and obstacle/danger-penalty handling were updated accordingly.Removed
Tree.backup_belief_v_from_children. Per-algorithm V-backup formulas are now inlined at the call sites — POMCP visited-only, iCVaR CVaR-over-children, othersmax.Removed the redundant
gammaparameter from Sparse PFT.PEP 639 license-file metadata now requires
setuptools>=77at build time.
New Features:
Added stochastic obstacle-collision and dangerous-area mechanics to
PushPOMDP,ContinuousPushPOMDP,RockSamplePOMDP,LaserTagPOMDPandContinuousLaserTagPOMDP. Bernoulli per-step penalty draws produce heavy-tailed return distributions for benchmarking risk-sensitive planners against expected-value MCTS on the same environment.Added a PacMan particle-belief visualization: a ghost-position particle overlay on the sprite viewer, with bundled DejaVu fonts for deterministic CI rendering.
Added typed accessors and compound mutation helpers to the arena tree (
increment_visit_count,update_action_q_with_return); all planners migrated to the typed surface.
Bug Fixes:
PFT-DPW / BetaZero: the immediate-reward stash is now keyed on
action_id; it was overwriting across sibling actions.POMCP-DPW: corrected saturated branch-reward propagation.
ConstrainedZero: aligned the failure target to a single episode-level target.
CVaR exploration: corrected the LCB formula, horizon-zero handling, LCB overflow on long horizons, and vectorized-belief support.
Sparse-sampling iCVaR: the unvisited-action mask was inverted.
Sparse sampling: fixed the branching-factor loop off-by-one.
BetaZero: unified continuous sampling across the rollout and tree-expansion paths.
PacMan: fixed multiple audit-flagged bugs — state encoding, reward sign on capture, terminal handling, and native/Python parity.
CartPole and ContinuousLightDark: observation-model corrections; dropped the ContinuousLightDark sampler grid-clip that biased particle weights.
LaserTag: fixed scalar/batch log-probability asymmetry, a pickling regression on the continuous variant, observation log-probability for the B1/B2 kernels, and a terminal-sentinel guard on
kernel.probability.Tiger: fixed listen-action impossible-observation handling; Push now advertises a reward range that includes the obstacle penalty.
RockSample: corrected the dangerous-area sign convention.
Continuous Push (discrete-actions variant): fixed obstacle-hit-probability forwarding.
Performance:
PFT-DPW belief sampling is amortized O(log K) via an inline CDF on weighted-particle beliefs.
Migrated the iCVaR CVaR-computation kernels to Numba; faster beacon-likelihood evaluation and systematic resampling on the iCVaR path.
Arena-tree column-store buffers are pre-sized to avoid reallocation during tree growth.
Added a Tiger-pattern sampling fast path to environment sampling.
Others:
Added iCVaR-POMCPOW tests pinning the LCB / CVaR-exploration formulas against the published reference.
Added arena-tree coverage and MCTS planner tree-structure tests.
Added env-API conformance tests for
hash_action/hash_observationacross all environments.Added a metric-invariants sanity suite — rate bounds, count non-negativity, CI bounds, return-shift linearity, belief invariants — wired into the per-environment metric tests.
Release 0.2.0 (2026-04-20)
Vectorized belief updaters and distributed-execution robustness
New Features:
Added vectorized belief updaters for the RockSample, PacMan, LightDark (continuous and discrete), CartPole, MountainCar, Push, Continuous Push, Continuous LaserTag and SafetyAnt environments, with batched NumPy updates for significant throughput gains.
Added observation-model-aware vectorized belief updaters for the LightDark family.
Added a
ParallelizationLeveloption for hyperparameter tuning, enabling episode-level parallelism alongside Optuna-level parallelism.Added Gaussian process noise to the CartPole and MountainCar state transition models.
Bug Fixes:
The CartPole and MountainCar vectorized updaters now correctly add process transition noise.
Removed duplicated reward logic and fixed an RNG-stream divergence in the
sample_next_steppaths.Post-run visualization no longer crashes Dask runs with
cannot pickle '_asyncio.Task'. Visualization is dispatched through the simulator’s task manager asEnvironmentVisualizationTasks, scaling across the full cluster instead of being capped by a local joblib pool.The distributed task pipeline is now OS-agnostic. Workers on a different OS than the client no longer die unpickling
pathlib.PosixPath:EnvironmentVisualizationTaskreturnsDict[str, bytes]from a worker-private scratch directory, and the episode / hyperparameter tuning tasks shipcache_dirasstrwith a graceful fallback to console-only logging when the path does not resolve on the worker’s OS.
Performance:
PushPOMDP.sample_next_step: ~5.2x speedup.RockSamplePOMDP.sample_next_step: ~6.35x speedup.DiscreteLightDarkPOMDP.sample_next_step: inlined sampling, pure-Python math for reward and beacon checks, squared-distance beacon proximity.DiscreteDistribution: faster initialization and sampling.CovarianceParameterizedMultivariateNormal: cached Cholesky transpose.
Others:
Added shared belief-level equivalence test utilities that validate vectorized updaters against non-vectorized baselines across environments.
Added a three-layer benchmark suite for planner and environment performance testing.
Added a weekly CI workflow running the full slow-test suite; 117 tests are marked
slow. The full suite also runs on pushes tomaster, while other branches and PRs skip slow tests.Split the Docker build into a reusable base image plus a thin CI layer, auto-building when the base image is missing from GHCR.
Added an auto-rebase workflow for open PRs whenever
developis updated.PacManPOMDPmethods now accept NumPy array states.Hyperparameter tuning computes
optuna_n_jobsandepisode_n_jobsonce in__init__.CartPole and MountainCar belief tests compare against deterministic physics rather than noisy samples.
Release 0.1.0 (2026-03-21)
Initial release
First public release of POMDPPlanners: core abstractions (Environment, Policy, Belief, Distributions), the MCTS planner family, the initial environment collection, and the simulation and hyperparameter-tuning framework.
Maintainers
POMDPPlanners is currently maintained by Yaacov Pariente (@yaacovpariente).