Optimizers¶
The optimizers submodule provides objects compatible with OpenQARP features that rely on the optimization of objective functions.
As the optimal strategy for circuit parameter optimization remains an open research question, this section outlines the usage of
available black-box optimizers and the process for implementing custom ones.
Unless otherwise specified, the Rosenbrock function will serve as the illustrative example throughout this section.
The corresponding Python implementation of the function and its gradient is as follows:
import numpy as np
def rosenbrock_function(x):
return np.sum(100 * (x[1:] - x[:-1]**2)**2 + (1 - x[:-1])**2)
def rosenbrock_gradient(x):
jac = np.zeros_like(x)
jac[:-1] += 200 * (x[1:] - x[:-1]**2) * (-2 * x[:-1]) - 2 * (1 - x[:-1])
jac[1:] += 200 * (x[1:] - x[:-1]**2)
return jac
Using the ScipyOptimizer¶
By leveraging SciPy for the minimization process, a wide range of solvers becomes accessible. Wrapping these solvers within our
own class ensures compatibility with other OpenQARP components. For a comprehensive list of available optimizers, refer to the
documentation for scipy.optimize.minimize.
As an example, we demonstrate the optimization of the Rosenbrock function using the BFGS quasi-Newton method.
To instantiate a ScipyOptimizer, specify the desired method and any options to be passed to the underlying SciPy optimizer
as a dictionary. The minimize method can then be invoked with the appropriate arguments:
from qarp.optimizers import ScipyOptimizer
optimizer = ScipyOptimizer(method="BFGS", options={"disp": True})
result = optimizer.minimize(
objective_function=rosenbrock_function,
initial_parameters=np.array([0., 0.]),
callback=None,
gradient=rosenbrock_gradient
)
That’s all that is required. Users are encouraged to consult the SciPy documentation for details on the options dictionary
and gradient support for each method.
Note
Spectrum estimation from DOS-QPE distributions (formerly the
SpectralOptimizer) is not an optimizer, since it implements no
minimize contract, and now lives in qarp.algorithms as
SpectrumEstimator. See the Spectrum estimation section of the
algorithms documentation.
RotosolveOptimizer¶
The RotosolveOptimizer implements the Rotosolve algorithm (Ostaszewski et. al. - Quantum 5, 391 (2021)), which analytically minimizes
each circuit parameter by exploiting the periodic structure of variational
quantum circuits.
Rotosolve assumes the objective function is periodic with period \(2\pi\) with respect to each parameter and updates parameters sequentially using three function evaluations per parameter.
The implementation supports optional learning‑rate scheduling and damped updates for improved convergence stability if needed.
Algorithm: Rotosolve
Input: Hermitian measurement operator \(M\) encoding the objective function, a parameterized quantum circuit \(U(\boldsymbol{\theta})\) with fixed structure, and a stopping criterion.
Initialize parameters \(\theta_d \in (-\pi, \pi]\) for \(d = 1, \ldots, D\) heuristically or at random.
Repeat for t:
For \(d = 1, \ldots, D\) do:
Select a value \(\phi \in \mathbb{R}\) heuristically or at random.
Fix all circuit parameters except the \(d\)‑th parameter.
Estimate the expectation values \(\langle M \rangle_{\phi}\), \(\langle M \rangle_{\phi + \frac{\pi}{2}}\), and \(\langle M \rangle_{\phi - \frac{\pi}{2}}\) from measurement samples.
Update the parameter:
\[\theta_d \leftarrow \phi - \frac{\pi}{2} - \operatorname{arctan2}\!\left( 2\langle M \rangle_{\phi} - \langle M \rangle_{\phi + \frac{\pi}{2}} - \langle M \rangle_{\phi - \frac{\pi}{2}}, \langle M \rangle_{\phi + \frac{\pi}{2}} - \langle M \rangle_{\phi - \frac{\pi}{2}} \right)\]
Until \(t=T_{max}\) or the stopping criterion is met.
Key Parameters:
maxiter: Maximum number of Rotosolve macro-iterationstol: Convergence tolerance on change in objective valuelr: Base learning rate applied to the analytic Rotosolve update (optional)schedule: Learning-rate schedule (optional)gamma: Exponential decay factor for learning-rate scheduling (optional)power: Power for polynomial learning-rate schedules (optional)verbose: Enable detailed macro- and micro-iteration logging
Note
The default Rotosolve does not have a learning rate scheduler, we include one as an heuristic option for smoother convergence according to a particular schedule with the update rule:
\[\theta_d \leftarrow \phi - \eta(t) \left( + \frac{\pi}{2} +\operatorname{arctan2}\!\left( 2\langle M \rangle_{\phi} - \langle M \rangle_{\phi + \frac{\pi}{2}} - \langle M \rangle_{\phi - \frac{\pi}{2}}, \langle M \rangle_{\phi + \frac{\pi}{2}} - \langle M \rangle_{\phi - \frac{\pi}{2}} \right)\right)\]
- \(\eta(t)\) is the learning rate computed via the schedule at iteration \(t\), and
if \(\eta(t) = 1 \forall t\) (which is the default), the algorithm defaults to vanilla Rotosolve.
Custom Optimizers¶
The structure of the optimizer classes allows for straightforward implementation of custom optimizers. To ensure compatibility with OpenQARP, custom optimizers must adhere to the following requirements:
The custom optimizer must implement a
.minimizemethod.The
.minimizemethod must accept the following arguments:objective_functionandinitial_parameters.The return value must be an object with an
.xattribute or property, consistent with SciPy’s behavior. No additional attributes are required for compatibility with other OpenQARP components.An optimizer that cannot honour
minimize(bounds=...)should set the class attributesupports_bounds = False(asGradientDescentOptimizerandRotosolveOptimizerdo) and raise on a non-Nonebounds; an algorithm whose search is bounded (MMQCELS) checks the attribute at construction and refuses such an optimizer with aTypeErrorinstead of failing mid-run. An object without the attribute is assumed to honour bounds.
Gradient-Based Optimizers¶
OpenQARP provides several built‑in gradient‑based optimizers implemented directly in
Python, using the ScipyOptimizer syntax including:
SGDOptimizer: Vanilla gradient descent/stochastic gradient descent optimizer.SPSAOptimizer: Optimiser for the SPSA algorithm, random vector perturbations with learning rate decay.RMSPropOptimizer: Gradient descent with adaptive squared‑gradient normalization.AdaGradOptimizer: Gradient descent with cumulative squared‑gradient scaling.AdamOptimizer: Gradient descent with Momentum and adaptive second-moment accumulation.AdamaxOptimizer: Infinity‑norm variant of Adam.NadamOptimizer: Nesterov‑accelerated Adam updates.
These optimizers follow a shared interface defined by
GradientDescentOptimizer, making them fully compatible with OpenQARP’s
variational algorithms and gradient approximation tools (finite differences fd,
SPSA-based (spsa) (random Rademacher perturbations) gradients, or analytic gradients if provided by the user).
Each optimizer implements only its update rule; all parameter validation, gradient construction, and input handling are centralized in the base class.
The fd and spsa estimators here call the objective once per point.
The engines expose the same estimators as batched gradient methods
(engine.run_gradient(params, method="finite-diff") /
method="spsa") that evaluate every point in one sweep, alongside the
analytic adjoint and parameter-shift methods — see Gradients.
Early Stopping¶
OpenQARP gradient-based optimizers provide an optional early‑stopping mechanism that can be enabled for any gradient‑based optimizer. Early stopping monitors the objective value during optimization and halts the optimization when the loss fails to improve by a minimum threshold for a specified number of consecutive steps. This helps avoid wasted function evaluations and prevents over‑optimization on noisy objectives.
The early‑stopping behaviour is controlled through the options dictionary
passed to the optimizer, alongside the usual keys such as maxiter:
from qarp.optimizers import SGDOptimizer
optimizer = SGDOptimizer(
options={
"maxiter": 200,
"early_stopping": True,
"es_tol": 1e-3,
"patience": 5,
"ema_beta": 0.9,
}
)
Parameters¶
early_stoppingbool, optionalEnables early stopping when set to
True. Default isFalse.es_tolfloat, optionalMinimum required improvement in the loss to be considered progress. If the smoothed loss (or raw loss, if EMA is disabled) does not improve by at least
es_tolrelative to the best observed value, the iteration is counted as a non‑improvement. Default is1e‑6.patienceint, optionalNumber of consecutive non‑improvement iterations tolerated before stopping. Once the number of consecutive failures to improve exceeds
patience, optimization terminates. Default is10.ema_betafloat, optionalExponential moving‑average smoothing factor. If set below
1.0, early stopping monitors the smoothed loss instead of the raw loss. A value such as0.9or0.99is typical for noisy objectives. Default is1.0, which disables smoothing.
Usage Notes¶
Early stopping works with all gradient estimation methods: analytic gradients, finite differences, and SPSA-computed (random directional derivative) gradients.
When
ema_betais set, smoothing can significantly stabilize stopping behaviour on stochastic or noisy objectives.Optimizers automatically return the best parameters encountered when early stopping is triggered, rather than those from the final iteration.
Since
ScipyOptimizerhas built in convergence halting routines for the optimisation, the EarlyStopping functionality mimics this forGradientDescentOptimizersubclasses, though it is not enabled by default.
SGDOptimizer¶
The SGDOptimizer implements the vanilla form of stochastic gradient descent
using a fixed (non-adaptive) learning rate. This optimizer represents the simplest gradient-based
method and serves as a strong baseline for smooth objective functions in both
classical and variational quantum optimization tasks.
How it works: At iteration \(k\), the update rule is:
where:
\(\eta\) is a constant learning rate,
\(\nabla f(\theta_k)\) is obtained analytically or approximately estimated.
SGD does not adjust \(\eta\) across iterations and does not maintain any per-parameter state such as momentum or variance estimates.
Key Options:
lr: Constant learning rate \(\eta\) (default: 0.01)grad_method:"fd"or"spsa"for gradient approximation.fd_eps: Finite-difference step used whengrad_method="fd"(default: 1e‑5)
Basic Usage:
from qarp.optimizers import SGDOptimizer
# Rosenbrock's curved valley is ill-conditioned for a fixed step size:
# lr much above 1e-3 here diverges. Vanilla SGD also needs more
# iterations than the adaptive methods below to make useful progress.
optimizer = SGDOptimizer(options={"lr": 0.001, "maxiter": 3000})
result = optimizer.minimize(
objective_function=rosenbrock_function,
initial_parameters=np.array([0., 0.])
)
print("Optimized parameters:", result.x)
Note
The “stochastic” in stochastic gradient descent typically refers to a gradient computed over a batched subset of input data in machine learning. When using this optimiser for variational quantum algorithms outside of quantum machine learning, there is typically no data. However, the SGD optimizer can still be stochastic due to measurement noise and gradient approximations.
SPSAOptimizer¶
The SPSAOptimizer implements the Simultaneous Perturbation Stochastic
Approximation (SPSA) algorithm, a highly efficient optimization
method that requires only two function evaluations per iteration, regardless
of the dimensionality of the parameter vector. Technically, SPSA computes directional derivatives in random
perturbed directions. This makes SPSA especially useful
for variational quantum algorithms (VQAs), noisy objective functions, and
high-dimensional optimization tasks.
Unlike standard finite-difference or quantum parameter-shift methods that scale linearly with the number of parameters, SPSA achieves dimensionality-independent cost by using random Rademacher perturbations to estimate gradients.
How it works: At iteration \(k\), SPSA samples a random perturbation vector \(\Delta_k\) whose entries take values \(+1\) or \(-1\) with equal probability. Two perturbed parameter vectors are evaluated:
The stochastic gradient estimate is:
The parameter update follows:
where \(a_k\) and \(c_k\) follow power-law decay schedules:
Typical recommended values (Spall, 1992) are:
\(\alpha \approx 0.602\)
\(\gamma \approx 0.101\)
Key Options:
spsa_a0: Initial learning-rate coefficient \(a_0\) (default: 0.1)spsa_alpha: Decay exponent \(\alpha\) (default: 0.602)spsa_A: Stabilizer parameter \(A\) in the \(a_k\) schedule (default: 10.0)spsa_c0: Initial perturbation scale \(c_0\) (default: 0.1)spsa_gamma: Perturbation decay exponent \(\gamma\) (default: 0.101)num_spsa: Number of SPSA perturbation pairs per iteration (default: 1)spsa_seed: Optional base random seed for reproducibility
Note
SPSA does not use analytical gradients. The grad_method argument is
automatically set to "spsa" when using SPSAOptimizer. This computes
the “gradients” as random directional derivatives multiplied by the perturbation.
Basic Usage:
from qarp.optimizers import SPSAOptimizer
# As with SGD, Rosenbrock's ill-conditioning means spsa_a0 much above
# 0.01 here diverges; more iterations make up for the smaller step.
optimizer = SPSAOptimizer(options={
"spsa_a0": 0.01,
"spsa_alpha": 0.602,
"spsa_A": 10.0,
"spsa_c0": 0.1,
"spsa_gamma": 0.101,
"spsa_seed": 123,
"maxiter": 3000,
})
result = optimizer.minimize(
objective_function=rosenbrock_function,
initial_parameters=np.array([0., 0.])
)
print("Optimized parameters:", result.x)
RMSPropOptimizer¶
The RMSPropOptimizer is an adaptive gradient method that normalizes parameter
updates by an exponential moving average of squared gradients. This stabilizes
optimization under noisy or highly anisotropic landscapes.
How it works: RMSProp maintains a running average:
and performs the update:
where:
\(\beta\) controls the decay of past gradients,
\(\varepsilon\) avoids division by zero.
Key Options:
lr: Learning rate \(\eta\) (default: 0.001)beta: Smoothing factor (default: 0.9)eps: Numerical stability term (default: 1e‑8)
Basic Usage:
from qarp.optimizers import RMSPropOptimizer
optimizer = RMSPropOptimizer(options={"lr": 0.001, "beta": 0.95})
result = optimizer.minimize(
objective_function=rosenbrock_function,
initial_parameters=np.array([-1.0, 1.0])
)
print("Optimized parameters:", result.x)
AdaGradOptimizer¶
The AdaGradOptimizer adapts the learning rate of each parameter individually
based on the accumulated squared gradients. This method is especially effective in
sparse or high‑dimensional optimization problems.
How it works: AdaGrad accumulates squared gradients:
and updates parameters according to:
As \(r_k\) grows, the effective learning rate decreases.
Key Options:
lr: Base learning rate \(\eta\) (default: 0.01)eps: Stability constant \(\varepsilon\) (default: 1e‑8)initial_accumulator_value: Starting value of the squared-gradient accumulator (default: 0.0)
Basic Usage:
from qarp.optimizers import AdaGradOptimizer
optimizer = AdaGradOptimizer(options={"lr": 0.02})
result = optimizer.minimize(
objective_function=rosenbrock_function,
initial_parameters=np.array([0., 0.])
)
print("Optimized parameters:", result.x)
AdamOptimizer¶
The AdamOptimizer combines momentum with adaptive per‑parameter learning
rates, making it one of the most widely used optimization algorithms in deep
learning and variational quantum eigensolvers.
How it works: Adam maintains exponential moving averages of gradients and squared gradients:
Bias-corrected estimates:
Update rule:
Key Options:
lr: Learning rate \(\eta\) (default: 0.001)beta1: Momentum decay (default: 0.9)beta2: Variance decay (default: 0.999)eps: Numerical stability constant (default: 1e‑10)
Basic Usage:
from qarp.optimizers import AdamOptimizer
optimizer = AdamOptimizer(options={"lr": 0.01})
result = optimizer.minimize(
objective_function=rosenbrock_function,
initial_parameters=np.array([-0.5, 0.5])
)
print("Optimized parameters:", result.x)
AdamaxOptimizer¶
The AdamaxOptimizer is an infinity-norm variant of Adam. By replacing the
second-moment estimate with a maximum-based accumulator, it improves robustness to
rare but very large gradient spikes.
How it works: Adamax keeps a momentum term:
and an exponentially weighted infinity norm:
Update rule:
Key Options:
lr: Learning rate \(\eta\) (default: 0.002)beta1: Momentum weight (default: 0.9)beta2: Infinity‑norm decay (default: 0.999)eps: Numerical stability constant (default: 1e‑8)
Basic Usage:
from qarp.optimizers import AdamaxOptimizer
optimizer = AdamaxOptimizer(options={"lr": 0.002})
result = optimizer.minimize(
objective_function=rosenbrock_function,
initial_parameters=np.array([1.0, 1.0])
)
print("Optimized parameters:", result.x)
NadamOptimizer¶
The NadamOptimizer introduces Nesterov momentum into Adam, producing a
look‑ahead‑corrected gradient estimate that often yields smoother convergence.
How it works: Momentum:
Second moment:
Nesterov‐corrected first moment:
Update rule:
Key Options:
lr: Learning rate \(\eta\) (default: 0.002)beta1: Momentum decay (default: 0.9)beta2: Variance decay (default: 0.999)eps: Numerical stability constant (default: 1e‑8)
Basic Usage:
from qarp.optimizers import NadamOptimizer
optimizer = NadamOptimizer(options={"lr": 0.002})
result = optimizer.minimize(
objective_function=rosenbrock_function,
initial_parameters=np.array([-1.0, 1.0])
)
print("Optimized parameters:", result.x)