Changelog

All notable changes to POMDPPlanners are documented here, newest release first. This project adheres to Semantic Versioning.

Each #NNN links to the pull request that introduced the change.

Release 0.5.0 (WIP)

CARLA, nuPlan and Isaac Lab environments; vectorized planning and batched GPU beliefs

Breaking Changes:

  • Environments now own their visualization file naming. Callers that built visualization paths themselves must go through the environment (#203).

New Features:

  • Added CarlaPOMDP, a forward-only world environment backed by the CARLA simulator (#204).

  • Added a pluggable generative-model interface and a multi-agent perception world for CARLA (#207).

  • Added planner-side perception in the CARLA world observation model, so the planner model and the world emit identical raw observations while the belief filters and infers hidden per-agent intent (#208).

  • Added a destination/route option for CARLA with a success terminal and route metrics (#221).

  • Added a headless CARLA server pool for parallel episode simulation (#217).

  • Added a nuPlan POMDP environment — world, model, perception and belief — with a POMCPOW example (#210).

  • Added an Isaac Lab POMDP environment and visualizer (#206).

  • Added a vectorized belief tree to POMDPPlanners.core (#214).

  • Added the VOPP vectorized planner and vectorized generative models (#215, #216).

  • Added BatchedParticleBelief, a torch-backed belief for batched GPU belief filtering (#218).

  • Added opt-in is_*_hit_terminal hazard-terminates-episode flags across the hazard environments, with a draw-coupled termination gate (#211, #212).

  • Added EVADE_WHEN_SPOTTED and PURSUE opponent-policy support to the vectorized LaserTag model (#219).

Bug Fixes:

  • Fixed nuPlan control commands, perception densities and metrics (#224).

  • Fixed the PacMan vectorized belief updater to model motion slip, which the scalar updater already applied (#213).

Documentation:

  • Rewrote the README to be user-focused (#223) and added CARLA and Isaac Sim images rendered by the package (#222).

  • Cited the Vectorized Online POMDP Planning paper on the VOPP planner (#225).

  • Removed the nuPlan evaluation notebook (#226).

Others:

  • Bumped the package version to 0.5.0 across pyproject.toml, POMDPPlanners/__init__.py, test_setup.py and the Sphinx release string, which had been stale at 0.2.0 (#228).

  • Refactored CARLA onto per-channel observation models with an encode_observation seam (#209).

  • Added coverage for the LaserTag belief updater honouring opponent_policy (#202) and for CARLA observation equality and collision-penalty reward (#205).

Release 0.4.0 (2026-05-28)

Constrained planners, hazard-centric reward variants, run-progress notifications

Breaking Changes:

  • Renamed the LightDark DANGEROUS_STATES reward model to HIGH_VARIANCE_STATES (#175).

  • Renamed reward variants to hazard-centric names across five environments (#186).

  • Dropped Python 3.8 and 3.9; the minimum supported version is now 3.10 (#192).

New Features:

  • Added the constrained planners CPOMCPOW and CPFT-DPW together with a ConstrainedEnvironment ABC (#170).

  • Added run-progress tracking: Slack notifications on run start, finish and failure, backed by a SQLite progress DB with an external stall-detection watcher for deaths the process cannot witness itself (#172).

  • Added per-task progress callbacks for the Dask and PBS backends (#173).

  • Added opt-in circular dangerous areas to PacManPOMDP (#176), plus a dangerous-area step counter and a total-danger-encounters aggregate (#187).

  • Added high-variance and decaying reward variants to RockSample — with batch acceleration — Push and PacMan (#180, #181, #182).

  • Added to_unique_support_distribution to VectorizedWeightedParticleBelief (#188).

  • Added a selectable LaserTag OpponentPolicy with evade, pursue and evade-when-spotted behaviours (#196, #197).

Bug Fixes:

  • Fixed the LaserTag opponent to evade and react to the robot’s pre-move position (#195).

  • Aligned the C++ reward kernels with their Python counterparts across five environments (#183) and closed Python-side reward-kernel bugs in PacMan and RockSample (#184).

  • Fixed three environment correctness bugs in the sanity, Tiger and LightDark environments (#185).

  • Fixed the PacMan visualizer to render the ghost belief heatmap for vectorized particle beliefs (#174).

Documentation:

  • Rewrote the README Architecture and Running Experiments sections (#190).

  • Standardized planner References: sections — one reference per planner, conference/published links preferred over arXiv (#200).

Others:

  • Replaced the deprecated pkg_resources with importlib.metadata (#171).

  • Extracted generic dangerous-area Numba kernels (#177) and reward models for RockSample, Push, LaserTag and PacMan (#178, #179).

  • Added MIT SPDX headers to all source files (#189).

  • The release workflow now builds an sdist only; the previous wheel carried an unportable linux_x86_64 platform tag that PyPI rejects (#193).

  • Made environment tests independent of constructor defaults (#199).

Release 0.3.1 (2026-05-12)

Packaging hotfix for a broken 0.3.0 sdist

Bug Fixes:

  • Bundled the C++ headers in the sdist. Without them pip install from PyPI failed to compile, because setup.py’s include_dirs pointed at paths absent from the tarball. Added packaging regression tests (#168).

Others:

  • Removed benchmark files from the repository root (#167).

  • Removed CHANGELOG.md (#166). Release notes returned as this page in 0.5.0.

Release 0.3.0 (2026-05-11)

Risk-sensitive benchmarking mechanics, arena tree, large performance and correctness pass

Breaking Changes:

  • Environment.reward gained an optional next_state parameter so penalty terms (obstacles, dangerous areas) are scored against the realised post-transition state rather than a fresh sample. Threaded through rollout and POMCP-DPW.

  • PacMan state is now a raw NumPy ndarray; the native batch path and obstacle/danger-penalty handling were updated accordingly.

  • Removed Tree.backup_belief_v_from_children. Per-algorithm V-backup formulas are now inlined at the call sites — POMCP visited-only, iCVaR CVaR-over-children, others max.

  • Removed the redundant gamma parameter from Sparse PFT.

  • PEP 639 license-file metadata now requires setuptools>=77 at build time.

New Features:

  • Added stochastic obstacle-collision and dangerous-area mechanics to PushPOMDP, ContinuousPushPOMDP, RockSamplePOMDP, LaserTagPOMDP and ContinuousLaserTagPOMDP. Bernoulli per-step penalty draws produce heavy-tailed return distributions for benchmarking risk-sensitive planners against expected-value MCTS on the same environment.

  • Added a PacMan particle-belief visualization: a ghost-position particle overlay on the sprite viewer, with bundled DejaVu fonts for deterministic CI rendering.

  • Added typed accessors and compound mutation helpers to the arena tree (increment_visit_count, update_action_q_with_return); all planners migrated to the typed surface.

Bug Fixes:

  • PFT-DPW / BetaZero: the immediate-reward stash is now keyed on action_id; it was overwriting across sibling actions.

  • POMCP-DPW: corrected saturated branch-reward propagation.

  • ConstrainedZero: aligned the failure target to a single episode-level target.

  • CVaR exploration: corrected the LCB formula, horizon-zero handling, LCB overflow on long horizons, and vectorized-belief support.

  • Sparse-sampling iCVaR: the unvisited-action mask was inverted.

  • Sparse sampling: fixed the branching-factor loop off-by-one.

  • BetaZero: unified continuous sampling across the rollout and tree-expansion paths.

  • PacMan: fixed multiple audit-flagged bugs — state encoding, reward sign on capture, terminal handling, and native/Python parity.

  • CartPole and ContinuousLightDark: observation-model corrections; dropped the ContinuousLightDark sampler grid-clip that biased particle weights.

  • LaserTag: fixed scalar/batch log-probability asymmetry, a pickling regression on the continuous variant, observation log-probability for the B1/B2 kernels, and a terminal-sentinel guard on kernel.probability.

  • Tiger: fixed listen-action impossible-observation handling; Push now advertises a reward range that includes the obstacle penalty.

  • RockSample: corrected the dangerous-area sign convention.

  • Continuous Push (discrete-actions variant): fixed obstacle-hit-probability forwarding.

Performance:

  • PFT-DPW belief sampling is amortized O(log K) via an inline CDF on weighted-particle beliefs.

  • Migrated the iCVaR CVaR-computation kernels to Numba; faster beacon-likelihood evaluation and systematic resampling on the iCVaR path.

  • Arena-tree column-store buffers are pre-sized to avoid reallocation during tree growth.

  • Added a Tiger-pattern sampling fast path to environment sampling.

Others:

  • Added iCVaR-POMCPOW tests pinning the LCB / CVaR-exploration formulas against the published reference.

  • Added arena-tree coverage and MCTS planner tree-structure tests.

  • Added env-API conformance tests for hash_action / hash_observation across all environments.

  • Added a metric-invariants sanity suite — rate bounds, count non-negativity, CI bounds, return-shift linearity, belief invariants — wired into the per-environment metric tests.

Release 0.2.0 (2026-04-20)

Vectorized belief updaters and distributed-execution robustness

New Features:

  • Added vectorized belief updaters for the RockSample, PacMan, LightDark (continuous and discrete), CartPole, MountainCar, Push, Continuous Push, Continuous LaserTag and SafetyAnt environments, with batched NumPy updates for significant throughput gains.

  • Added observation-model-aware vectorized belief updaters for the LightDark family.

  • Added a ParallelizationLevel option for hyperparameter tuning, enabling episode-level parallelism alongside Optuna-level parallelism.

  • Added Gaussian process noise to the CartPole and MountainCar state transition models.

Bug Fixes:

  • The CartPole and MountainCar vectorized updaters now correctly add process transition noise.

  • Removed duplicated reward logic and fixed an RNG-stream divergence in the sample_next_step paths.

  • Post-run visualization no longer crashes Dask runs with cannot pickle '_asyncio.Task'. Visualization is dispatched through the simulator’s task manager as EnvironmentVisualizationTasks, scaling across the full cluster instead of being capped by a local joblib pool.

  • The distributed task pipeline is now OS-agnostic. Workers on a different OS than the client no longer die unpickling pathlib.PosixPath: EnvironmentVisualizationTask returns Dict[str, bytes] from a worker-private scratch directory, and the episode / hyperparameter tuning tasks ship cache_dir as str with a graceful fallback to console-only logging when the path does not resolve on the worker’s OS.

Performance:

  • PushPOMDP.sample_next_step: ~5.2x speedup.

  • RockSamplePOMDP.sample_next_step: ~6.35x speedup.

  • DiscreteLightDarkPOMDP.sample_next_step: inlined sampling, pure-Python math for reward and beacon checks, squared-distance beacon proximity.

  • DiscreteDistribution: faster initialization and sampling.

  • CovarianceParameterizedMultivariateNormal: cached Cholesky transpose.

Others:

  • Added shared belief-level equivalence test utilities that validate vectorized updaters against non-vectorized baselines across environments.

  • Added a three-layer benchmark suite for planner and environment performance testing.

  • Added a weekly CI workflow running the full slow-test suite; 117 tests are marked slow. The full suite also runs on pushes to master, while other branches and PRs skip slow tests.

  • Split the Docker build into a reusable base image plus a thin CI layer, auto-building when the base image is missing from GHCR.

  • Added an auto-rebase workflow for open PRs whenever develop is updated.

  • PacManPOMDP methods now accept NumPy array states.

  • Hyperparameter tuning computes optuna_n_jobs and episode_n_jobs once in __init__.

  • CartPole and MountainCar belief tests compare against deterministic physics rather than noisy samples.

Release 0.1.0 (2026-03-21)

Initial release

  • First public release of POMDPPlanners: core abstractions (Environment, Policy, Belief, Distributions), the MCTS planner family, the initial environment collection, and the simulation and hyperparameter-tuning framework.

Maintainers

POMDPPlanners is currently maintained by Yaacov Pariente (@yaacovpariente).