LINE Solver (C++)
Templated C++ port of the LINE queueing solver
Loading...
Searching...
No Matches
solver_ln.h
Go to the documentation of this file.
1/*
2 * Copyright (c) 2012-2026, QORE Lab, Imperial College London
3 * All rights reserved.
4 */
5#ifndef LINE_SOLVERS_LN_SOLVER_LN_H
6#define LINE_SOLVERS_LN_SOLVER_LN_H
7
8/**
9 * @file
10 * @ingroup line_solvers
11 * SolverLN: layered decomposition of a layered queueing network.
12 *
13 * Port of the matlab/src/solvers/LN SolverLN class folder -- construct, buildLayers,
14 * buildLayersRecursive, init, initInterlock, converged, analyze, post,
15 * updateMetrics, updatePopulations, updateThinkTimes, updateLayers,
16 * updateRoutingProbabilities, getEntryServiceMatrix and getEnsembleAvg --
17 * driven by the EnsembleSolver iteration in @@EnsembleSolver/iterate.m.
18 *
19 * The method is a Picard iteration over a fixed point. Each processor and each
20 * software task becomes a closed queueing network (a "layer") in which the
21 * element is the server and its callers are the customers; solving all layers
22 * gives new residence times, which become the service times of the calls one
23 * level up and the think times of the callers one level down, and the layers
24 * are re-solved until the queue lengths stop moving.
25 *
26 * WHAT IS IMPLEMENTED. Every path a plain layered model takes is here:
27 * synchronous calls, entry selection by throughput ratio, multi-entry tasks,
28 * reference tasks with think time, infinite-server tasks and processors,
29 * server replication, AND-forks and AND-joins, and the LQNS V5 interlock
30 * analysis. Forwarding is solved too: construct() runs the
31 * lqn_fwd_rendezvous rewrite of lqn_helpers.h (:306), which flattens every
32 * forwarding chain into caller-side pseudo rendezvous arcs, so nothing below
33 * this point needs a forwarding case. Asynchronous calls and entry open
34 * arrivals are solved by giving the layer its own Source/Sink pair and an
35 * open chain (build_layer's async_here/open_entries construction, roughly
36 * :622-668). Cache tasks and item entries are solved by a Cache node placed
37 * in the host layer, with the hit and miss branches switching class at the
38 * server (build_layer's iscachelayer path, :591-616 and the read walk at
39 * :883-908). Setup tasks with a setup time are solved by charging the cold
40 * start to the ENTRY, at the probability the thread was found powered down
41 * (setup_charge, added to entry_servt and to the caller's think time); no layer
42 * station carries a SetupDelayOffParam. Admission constraints are solved by a Region on
43 * the server station, which forces that layer's fixed point through
44 * SolverCTMC rather than MVA or fluid (:1051-1096, dispatch at :1753-1760).
45 * reject_unsupported() (:433) is kept but empty: every construct the
46 * reference LQN model can carry is now accounted for here. A few narrower
47 * combinations remain unsupported and are refused where they are discovered
48 * instead of up front: a fork sharing a layer with an open stream (:492-496),
49 * a fork's fixed point outside layer_solver 'mva' (:1723-1728), and a fluid
50 * layer or an AND-join's order-statistic fit under a non-double/non-real
51 * arithmetic backend (:120-126, :2198-2201).
52 *
53 * WHAT ELSE THIS CLASS ANSWERS, beyond the mean-value table. `options.method`
54 * selects between three DIFFERENT questions rather than three routes to one:
55 * `default` is the mean-based update; `moment3` is updateMetricsMomentBased,
56 * which fits an APH to each layer's response-time CDF and convolves along the
57 * entry's activity sequence, so it additionally reports a per-entry response
58 * time DISTRIBUTION through get_cdf_respt(); `mw.upper` / `mw.lower` report
59 * Majumdar-Woodside robust box bounds and never solve a layer at all.
60 * get_tran_avg() is the layered transient, decoupled (demands frozen at the
61 * fixed point) or coupled (waveform relaxation over time-varying demands
62 * injected through the fluid rate schedule). get_sensitivity_table() delegates
63 * to each layer's own sensitivity table; every derivative in it is a partial
64 * WITHIN its layer, not a total derivative of the layered model.
65 *
66 * THE LAYER ENGINE is `layer_solver`: `mva`, `nc`, `fluid` or `ssa`, the C++
67 * spelling of the reference's solver FACTORY. They converge to DIFFERENT fixed
68 * points, because each feeds different demands back into the next outer sweep.
69 * `ssa` is noisy, so the deterministic convergence test is replaced by
70 * LnStochController (Robbins-Monro relaxation, Polyak-Ruppert reporting);
71 * running it under the deterministic test would simply reach iter_max.
72 *
73 * PHASE-2 ACTIVITIES ARE SOLVED, not refused. `servt` keeps both phases, so the
74 * server's utilization is unchanged, while `residt` carries the CALLER's view:
75 * phase 1 in full plus the phase-2 time the caller is actually overtaken by,
76 * with the overtaking probability from lqn_analyzers.h's
77 * lqn_overtake_prob_markov. Every phase-2 branch below is gated on `has_phase2`
78 * and is inert on a single-phase model.
79 *
80 * ARITHMETIC. All model quantities are T. Iteration counters, populations,
81 * multiplicities and the convergence tolerances are double, matching the
82 * reference: they are properties of the algorithm and of the model's integer
83 * structure, not quantities whose precision is under study.
84 */
85
87#include <algorithm>
88#include <cctype>
89#include <cmath>
90#include <functional>
91#include <iostream>
92#include <limits>
93#include <map>
94#include <memory>
95#include <set>
96#include <string>
97#include <vector>
98
104#include "line/api/mam/aph_fit.h"
116#include "line/api/lqn/lqn_ph.h"
126#include "line/util/error.h"
127#include "line/util/expm.h"
128#include "line/util/linalg.h"
129#include "line/util/matrix.h"
130
131namespace line {
132namespace ln {
133
134using lang::CallType;
135using lang::Distrib;
136using lang::GlobalConstants;
138using lang::LqnElement;
139using lang::NodeType;
143using lqn::LqnCallGroup;
144using lqn::LqnStruct;
145
146/** One (station, class) rate trajectory injected into a layer's closing ODE. */
148
149namespace detail {
150
151/**
152 * Run the fluid analyzer on one layer and write its metrics into the shape the
153 * MVA layer path returns.
154 *
155 * Overloaded rather than gated: LSODA is double-only, so the template below is
156 * what any other backend resolves to, and it refuses by name instead of failing
157 * to compile inside a branch that could never run.
158 */
159inline void ln_fluid_solve(const qn::NetworkStruct<double>& L, const fluid::FluidOptions& fo,
161 // THE RUNNER, NOT THE SWITCH. `solver_fluid` is the ungated port of
162 // solver_fluid_analyzer.m and cannot reach `minnormal`, `rmf`, `refined` or
163 // `kp` at all; `solver_fluid_run_analyzer` is the port of runAnalyzer's resolution,
164 // which is what `LN(model, @(m) Fluid(m))` runs in the reference. Calling
165 // the switch here silently ran the FIRST-ORDER `matrix` method in every
166 // layer: on lqn_basic that reported T3's residence as its bare demand
167 // (0.02 against MATLAB's 0.0206615) because the hard min charges no
168 // queueing below the server count, and the LN table then disagreed with
169 // MATLAB by 3.3% on a metric neither codebase flagged.
171 for (std::size_t i = 0; i < L.nstations; ++i)
172 for (std::size_t r = 0; r < L.nclasses; ++r) {
173 out.Q(i, r) = s.QN(i, r);
174 out.U(i, r) = s.UN(i, r);
175 out.R(i, r) = s.RN(i, r);
176 out.Tp(i, r) = s.TN(i, r);
177 }
178 for (std::size_t r = 0; r < L.nclasses && r < s.XN.size(); ++r) {
179 out.X[r] = s.XN[r];
180 out.C[r] = s.CN[r];
181 }
182 out.method = s.method;
183 out.iter = static_cast<int>(s.iters);
184}
185template <class T>
186void ln_fluid_solve(const qn::NetworkStruct<T>&, const fluid::FluidOptions&,
188 throw UnsupportedError(
189 "SolverLN: a fluid layer integrates its drift with LSODA, whose coefficients assume "
190 "double precision; rerun with --arith double or use layer_solver 'mva'");
191}
192
193/**
194 * The response-time CDF of every (station, class) of one layer, which is what
195 * `SolverFluid(ensemble{e}).getCdfRespT` returns in the reference.
196 *
197 * The layer is integrated once to its fixed point and the converged state is
198 * the marking the passage starts from, so the law reported is the STATIONARY
199 * response time and not the one seen from an arbitrary initial condition. That
200 * is `options.init_sol = odeStateVec` in `@@SolverFLD/getCdfRespT`.
201 *
202 * A pair the class never visits comes back as the degenerate curve at zero,
203 * which `fluid_passage_time` already returns; the moment extraction above reads
204 * it as a mean of zero and the caller drops the term, as MATLAB's
205 * `m1 > CoarseTol` test does.
206 */
207inline std::vector<std::vector<fluid::FluidPassage>> ln_fluid_cdf_respt(
209 fluid::FluidOptions o = fo;
210 // THE SWITCH HERE, DELIBERATELY, unlike ln_fluid_solve above. By the same
211 // argument this should be the runner, but the reference for `moment3` on
212 // this path is the JAR (MATLAB does not finish the method on the model the
213 // golden was taken from), and the JAR's entry response times agree with the
214 // first-order switch to 5e-3 and not with the resolved method, which moves
215 // E:e0 from 2.2748 to 2.4964. Switching it blind would replace a measured
216 // agreement with an unmeasured one; it needs a JAR-side reading first.
218 std::vector<std::vector<fluid::FluidPassage>> out(
219 L.nstations, std::vector<fluid::FluidPassage>(L.nclasses));
220 for (std::size_t i = 1; i <= L.nstations; ++i) {
221 if (L.stations[i - 1].nodetype == qn::NodeType::Source) continue;
222 for (std::size_t r = 1; r <= L.nclasses; ++r) {
223 if (L.disabled[i - 1][r - 1]) continue;
224 out[i - 1][r - 1] = fluid::fluid_passage_time(L, s.xvec, i, r, o.tol, 201, s.closure);
225 }
226 }
227 return out;
228}
229template <class T>
230std::vector<std::vector<fluid::FluidPassage>> ln_fluid_cdf_respt(const qn::NetworkStruct<T>&,
231 const fluid::FluidOptions&) {
232 throw UnsupportedError(
233 "SolverLN: the response-time CDF a layer contributes to the 'moment3' update is a fluid "
234 "passage time, integrated by LSODA in double precision; rerun with --arith double");
235}
236
237/**
238 * One layer's transient trajectory, the `getTranAvg` of a layer.
239 *
240 * `out_grid` REPLACES the uniform grid, and it is not a convenience: SolverENV
241 * sums a layered stage's exit average against the environment's holding-time
242 * CDF, so the trajectory has to be evaluated where that CDF puts its mass
243 * rather than where the horizon happens to spread points. It also makes every
244 * layer of one stage report on the SAME grid, which is what lets the layers be
245 * assembled into one block-diagonal transient with a single time base.
246 */
247inline std::vector<fluid::FluidTranPoint> ln_fluid_transient(
248 const qn::NetworkStruct<double>& L, const fluid::FluidOptions& fo, double t_end,
249 std::size_t points, const std::vector<double>& out_grid = std::vector<double>()) {
250 return fluid::solver_fluid_transient(L, fo, t_end, points, out_grid);
251}
252template <class T>
253std::vector<fluid::FluidTranPoint> ln_fluid_transient(const qn::NetworkStruct<T>&,
254 const fluid::FluidOptions&, double,
255 std::size_t,
256 const std::vector<double>& =
257 std::vector<double>()) {
258 throw UnsupportedError(
259 "SolverLN: the layered transient integrates each layer's drift with LSODA, which is "
260 "double precision by construction; rerun with --arith double");
261}
262
263} // namespace detail
264
265/** Options of SolverLN. Defaults are SolverOptions('LN'). */
266struct LnOptions {
267 int iter_max = 200;
268 double iter_tol = 5e-3;
269 double tol = 1e-4;
270 bool interlocking = true;
271 /**
272 * `options.config.interlock_method`: how the interlock is TRACKED once
273 * `interlocking` is on. `ilrate` is the post-hoc rate discount of
274 * update_populations (the default and the historical method); `refpath`
275 * merges the callers that are the same reference customers into ONE client
276 * chain running down the reference path, replacing the discount; `none`
277 * applies no correction. A request refpath cannot carry DOWNGRADES to
278 * `ilrate`, see ln_interlock_method.
279 */
280 std::string interlock_method = "ilrate";
281 /** `options.config.interlock_maxpaths`: refpath refuses a layer carrying more routes into it. */
282 double interlock_maxpaths = 32.0;
283 /**
284 * `options.config.interlock_refpath_scope`: `merging` transforms only the
285 * layers where two or more callers descend from one reference task; `all`
286 * also those with a lone caller below a reference task.
287 */
288 std::string interlock_refpath_scope = "merging";
289 std::string relax = "fixed";
290 double relax_factor = 0.5;
291 /** Options handed to each layer solver; SolverMVA defaults. */
293 /**
294 * Which solver runs each layer: `mva`, `nc`, `fluid` or `ssa`.
295 *
296 * The reference names this by passing a solver FACTORY,
297 * `LN(model, @(m) Fluid(m, opt))`, so the choice is per-ensemble and not
298 * per-layer; this field is the same choice under a name, because a C++
299 * template cannot take a MATLAB function handle.
300 *
301 * THEY DO NOT CONVERGE TO THE SAME PLACE. Each engine feeds different
302 * service demands back into the next outer iteration, so the ensembles
303 * reach different fixed points rather than the same one by different
304 * routes. `ssa` in particular is a NOISY layer solver: the deterministic
305 * convergence test cannot terminate against its standard error, so it is
306 * driven with `LnStochController` (lqn_analyzers.h) instead.
307 */
308 std::string layer_solver = "mva";
309 /** Options handed to each layer when `layer_solver` is `fluid`. */
311 /** Options handed to each layer when `layer_solver` is `nc`. */
313 /** Options handed to each layer when `layer_solver` is `ssa`. */
315 /**
316 * `options.method`, which selects WHAT is reported and not merely how:
317 *
318 * `default` the mean-based update (updateMetricsDefault)
319 * `moment3` the APH moment-based update (updateMetricsMomentBased),
320 * which additionally produces a per-entry response-time
321 * distribution, `get_cdf_respt()`
322 * `mw.upper` Majumdar-Woodside robust box bounds INSTEAD of a fixed
323 * `mw.lower` point: no layer is ever solved, and every metric the
324 * bound does not define is reported as undefined
325 */
326 std::string method = "default";
327 /**
328 * `options.config.ln_transient`: how the per-layer transients are coupled.
329 * `decoupled` freezes the inter-layer demands at the converged fixed point;
330 * `coupled` reconciles them by waveform relaxation. Iteration 0 of the
331 * coupled relaxation IS the decoupled answer.
332 */
333 std::string ln_transient = "coupled";
334 /** `options.config.ln_transient_iter_max` and `..._tol` of the relaxation. */
336 double ln_transient_tol = 1e-2;
337 /**
338 * `options.config.ln_transient_channels`: which inter-layer coupling is
339 * injected, `both`, `thinkt` (client delay only) or `callservt`
340 * (synchronous-call service only). Used to isolate each channel's share of
341 * the coupled transient.
342 */
343 std::string ln_transient_channels = "both";
344 /** `options.timespan(2)`: the transient horizon; infinite means none is set. */
345 double timespan_end = std::numeric_limits<double>::infinity();
346 /** Output points per layer trajectory, and the relaxation's own grid size. */
347 std::size_t tran_points = 101;
348 /**
349 * An EXPLICIT output grid for the layered transient, replacing the uniform
350 * `tran_points` one when it is not empty.
351 *
352 * SolverENV is what asks for it: a layered stage's exit metric is a
353 * Riemann-Stieltjes sum against the environment's holding-time CDF, and a
354 * uniform grid over the horizon resolves the horizon rather than the
355 * sojourn. Interpolating afterwards cannot recover resolution the
356 * trajectory never had, so the points are asked for instead --
357 * `SolverEnv::stage_grid` builds them and `set_tran_grid` installs them.
358 * The grid must be increasing and start at zero; LSODA takes it as given.
359 */
360 std::vector<double> tran_grid;
361};
362
363/** Per-layer results of one iteration, the [QN,UN,RN,TN,AN,WN] of getAvg. */
364template <class T>
367};
368
369/** The LQN-level answer, indexed by element 1..nidx. */
370template <class T>
372 std::vector<T> QN, UN, RN, TN, AN, WN;
374 int iterations = 0;
375 bool converged = false;
376 /**
377 * True when the numbers are a BOUND (`method` = `mw.upper` / `mw.lower`)
378 * rather than the fixed point. A bound defines throughput and processor
379 * utilization and nothing else, so the other four measures come back
380 * undefined; reporting zeros there would be a claim.
381 */
382 bool is_bound = false;
383};
384
385/** A CDF sampled on a grid, the [F, t] pair MATLAB's evalCDF returns. */
386struct LnCdf {
387 std::vector<double> t, cdf;
388 bool empty() const { return t.empty(); }
389};
390
391/**
392 * One layer's block of the layered transient.
393 *
394 * The reference assembles the layers into one BLOCK-DIAGONAL cell array whose
395 * off-diagonal blocks are empty, so the blocks themselves carry the whole
396 * answer; keeping them apart also keeps each layer's own time grid, which the
397 * block-diagonal form has nowhere to put.
398 */
400 std::vector<double> t; ///< the output grid, shared by every series below
401 /// [station][class][point]
402 std::vector<std::vector<std::vector<double>>> QN, UN, TN;
403};
404
405/**
406 * `LayeredNetwork.layerBlocks`: where each layer's block sits in the aggregate.
407 *
408 * An LQN has no stations and classes of its own -- SolverLN builds them, one
409 * network per layer -- so the flat (station, class) view a caller like
410 * SolverENV needs is the BLOCK-DIAGONAL UNION of the layer networks, in layer
411 * order. `roff[e]`/`coff[e]` are layer e's 0-based row/column offset in that
412 * union and `msz[e]`/`ksz[e]` its size; `M`/`K` are the totals.
413 *
414 * The off-diagonal blocks pair a station of one layer with a class of another
415 * and stand for nothing at all. The reference leaves them as EMPTY cells; this
416 * port fills them with zeros, which is what the reference's consumers make of
417 * an empty cell (`tranTimeBase_` returns nothing and the metric stays 0).
418 */
420 std::vector<std::size_t> roff, coff, msz, ksz;
421 std::size_t M = 0, K = 0;
422};
423
424/** The layered transient: one block per layer, plus how it was produced. */
426 std::vector<LnTranLayer> layers;
427 std::string mode; ///< "coupled" or "decoupled"
428 long iterations = 0; ///< waveform-relaxation sweeps; 0 when decoupled
429 double gap = 0.0; ///< final sup-norm trajectory change, coupled only
430};
431
432/** getSensitivityTable of the ensemble: the layer tables under a Layer column. */
433template <class T>
435 struct Row {
436 std::string layer, station, jobclass;
438 };
439 std::vector<Row> rows;
440 /** Per layer, the branch that layer took; empty for a layer with no solver. */
441 std::vector<std::string> layer_methods;
442 /** The summary label: the common branch, or "mixed" when they differ. */
443 std::string method;
444 /** Per layer, the analytic Jacobian where that layer produced one. */
445 std::vector<sens::SensTable<T>> layer_tables;
446};
447
448/**
449 * Overtaking probability at a server entry, defined in lqn_analyzers.h.
450 *
451 * DECLARED, not included: lqn_analyzers.h needs LayerResult and SolverLN
452 * complete (LnStochController holds a vector of the first and two refusals take
453 * the second), so it must be parsed AFTER this class. The definition arrives
454 * through the include at the foot of this file, which is why the declaration
455 * has to stand here -- an unqualified call from a member function would
456 * otherwise find nothing, LqnStruct's associated namespace being line::lqn.
457 */
458template <class T>
459T lqn_overtake_prob_markov(const LqnStruct<T>& lqn, const std::vector<T>& servt,
460 const std::vector<T>& callresidt, const std::vector<T>& tput,
461 std::size_t eidx, const T& xj);
462
463/** `options.config.stochiter_*` of SolverOptions.m, with its defaults. */
465 long burnin = 5; ///< Picard iterations before the step decay starts
466 double a0 = 1.0; ///< Robbins-Monro step immediately after burn-in
467 double alpha = 0.6; ///< step decay exponent, in (0.5, 1]
468 long conseq = 3; ///< consecutive sub-tolerance iterations required to stop
469 double iter_tol = 5e-3;
470 /** Relaxation in force during burn-in, i.e. whatever init left in place. */
471 double relax_burnin = 1.0;
472};
473
474/**
475 * The Robbins-Monro / Polyak-Ruppert controller, defined in lqn_analyzers.h.
476 *
477 * DECLARED HERE, not included: same mutual dependence as above. `iterate` holds
478 * one through a shared_ptr rather than by value so that this class stays
479 * complete without it -- a by-value member would need the definition at the
480 * point the class template is instantiated, and the definition arrives at the
481 * foot of this file.
482 *
483 * ITS CONFIG CANNOT BE DEFERRED THE SAME WAY. `iterate` names LnStochConfig by
484 * value, and that name does not depend on T, so it is looked up and required
485 * COMPLETE when the template is parsed rather than when it is instantiated --
486 * the shared_ptr trick only defers the class template beside it.
487 */
488template <class T>
490
491template <class T>
492class SolverLN {
493public:
494 /**
495 * Port of `SolverLN.listValidMethods`.
496 *
497 * Each name states the LAYERING and the ENCODING; `ln_requested_method`
498 * normalises the alias spellings ("ph", "cs", "srvncs", "flatcs",
499 * "squashed", "squashed.ph") onto these, and they are left out here to keep
500 * the list unambiguous, exactly as the reference does.
501 */
502 static std::vector<std::string> list_valid_methods() {
503 return {"srvn", "srvn.ph", "srvn.cs", "flat", "flat.cs", "flat.ph", "moment3", "default"};
504 }
505
506 SolverLN(const LqnStruct<T>& lqn_in, const LnOptions& options) : lqn(lqn_in), opt(options) {
507 // A LAYER SOLVER NEVER NARRATES, the same invariant the MATLAB, JAR and
508 // python ports stamp on each constructed layer solver. The fixed point
509 // runs every layer once per iteration, so a layer left at the caller's
510 // verbosity would print its own banner nlayers*iter_max times and bury
511 // the layered narration the caller actually asked for. `ssa` is the one
512 // layer engine here carrying a verbosity knob of its own; the level is
513 // forced rather than trusted so a library caller cannot set it.
514 // SolverLN's own reporting is unaffected.
515 opt.layer_ssa.verbose = false;
516 // interlock_method itself is validated by ln_interlock_method, where it is resolved
517 for (char& ch : opt.interlock_refpath_scope)
518 ch = char(std::tolower(static_cast<unsigned char>(ch)));
519 if (opt.interlock_refpath_scope != "merging" && opt.interlock_refpath_scope != "all")
520 throw InputError("Unknown config.interlock_refpath_scope '" +
521 opt.interlock_refpath_scope + "', use 'merging' or 'all'.");
522 if (std::isnan(opt.interlock_maxpaths) || opt.interlock_maxpaths < 0.0)
523 throw InputError("config.interlock_maxpaths must be a non-negative number.");
524 construct();
525 }
526
527 /** Port of getEnsembleAvg: run the iteration and aggregate onto LQN elements. */
529 // getAvgTable.m:611-624 answers the two bound requests WITHOUT running
530 // the fixed point at all: a box bound is a statement about the model,
531 // not about an iterate, so solving first and discarding the answer would
532 // only cost time and invite the two to be confused.
533 if (opt.method == "mw.upper" || opt.method == "mw.lower") return box_bounds();
534 iterate();
535 // the layers of "srvn.ph" carry one class per caller task, so the
536 // per-element results are rebuilt analytically -- see aggregate_ph
537 return is_ph_encoding() ? aggregate_ph() : aggregate();
538 }
539
540 /**
541 * Port of @@SolverLN/getCdfRespT: the per-entry response-time distribution.
542 *
543 * ONLY THE `moment3` METHOD PRODUCES ONE. The reference reacts to an empty
544 * table by re-running getAvg under method='moment3' and restoring the
545 * caller's method afterwards, and so does this: the distribution is the
546 * whole point of that method, and the mean-based update has no distribution
547 * to report, not a coarser one.
548 *
549 * Indexed by entry NUMBER 1..nentries, as the reference's
550 * `entrycdfrespt{eidx - (nhosts+ntasks)}` is.
551 */
552 std::vector<LnCdf> get_cdf_respt() {
553 if (lqn.nentries == 0) return entrycdfrespt;
554 if (entrycdfrespt.size() <= lqn.nentries || entrycdfrespt[1].empty()) {
555 // The distribution pass reads the ROUTING encoding of the activity
556 // graph, which srvn.ph layers do not carry: re-running the iteration
557 // over them would reconstruct the wrong topology rather than a
558 // coarser answer, and returning the empty table would report no
559 // distribution at all. Refuse by name, as the reference does
560 // (@SolverLN/getCdfRespT).
561 if (is_ph_encoding())
562 throw UnsupportedError(
563 "getCdfRespT needs the routing encoding of the activity graph, which "
564 "method='srvn.ph' does not build. Rebuild the solver with "
565 "method='srvn.cs' or method='moment3'.");
566 // BOTH the option and the RESOLVED method have to move. update_metrics
567 // dispatches on `lnmethod`, which build_layers resolved once, so
568 // flipping `opt.method` alone leaves the mean-based update in place and
569 // the table empty -- which is what this getter did between the alias
570 // landing and 2026-08-11. The routing layers already built serve
571 // moment3 unchanged, so only the update pass changes.
572 const std::string saved = opt.method;
573 const std::string saved_resolved = lnmethod;
574 opt.method = "moment3";
575 lnmethod = "moment3";
576 iterate();
577 aggregate();
578 opt.method = saved;
579 lnmethod = saved_resolved;
580 }
581 return entrycdfrespt;
582 }
583
584 /**
585 * Port of @@SolverLN/getTranAvg: the block-diagonal aggregate transient.
586 *
587 * `opt.ln_transient` selects the coupling; both modes return the same
588 * layout, and iteration 0 of the coupled relaxation IS the decoupled
589 * answer.
590 */
592 if (opt.ln_transient == "decoupled") return tran_avg_decoupled();
593 if (opt.ln_transient == "coupled") return tran_avg_coupled();
594 throw InputError("SolverLN: unknown ln_transient mode '" + opt.ln_transient +
595 "' (use 'coupled' or 'decoupled')");
596 }
597
598 /**
599 * Port of @@SolverLN/getSensitivityTable: solve the ensemble, then
600 * concatenate each layer solver's own table under a leading Layer column.
601 *
602 * WHAT THESE DERIVATIVES MEAN, and the reference is emphatic about it: each
603 * entry is a partial derivative WITHIN ITS LAYER, taken with the layer
604 * parameters the fixed point produced held fixed. Perturbing a host demand
605 * moves the think times, populations and demands of every other layer
606 * through the fixed-point map, and that indirect term is NOT included. The
607 * table attributes a bottleneck inside a layer; it does not predict the
608 * effect of a parameter change on the solved layered model.
609 */
611 if (results.empty()) iterate();
612 LnSensTable<T> out;
613 out.layer_methods.assign(ensemble.size(), std::string());
614 out.layer_tables.resize(ensemble.size());
615 for (std::size_t e = 0; e < ensemble.size(); ++e) {
616 sens::SensOptions so = sopt;
617 so.simulation = opt.layer_solver == "ssa";
618 const bool exact_available =
619 opt.layer_solver == "mva" || opt.layer_solver == "nc";
621 ensemble[e], so, exact_available && !fj_tr[e].active(),
622 [this, e]() { return this->solve_layer(e); });
623 out.layer_methods[e] = t.method;
624 for (const sens::SensRow<T>& r : t.rows) {
625 typename LnSensTable<T>::Row row;
626 row.layer = ensemble[e].name;
627 row.station = r.station;
628 row.jobclass = r.jobclass;
629 row.dTput = r.dTput;
630 row.dRespT = r.dRespT;
631 row.dQLen = r.dQLen;
632 row.dUtil = r.dUtil;
633 out.rows.push_back(row);
634 }
635 out.layer_tables[e] = t;
636 }
637 // One branch label per layer, plus a summary that is `mixed` when the
638 // layers did not all take the same branch.
639 for (const std::string& m : out.layer_methods) {
640 if (m.empty()) continue;
641 if (out.method.empty()) out.method = m;
642 else if (out.method != m) out.method = "mixed";
643 }
644 return out;
645 }
646
647 /**
648 * Majumdar-Woodside robust box bounds, reported in the shape of a solution.
649 *
650 * `mw.upper` reports the upper bound of both throughput and processor
651 * utilization, `mw.lower` the lower bound of both. Everything the bound
652 * does not define -- queue length, response time, residence time, arrival
653 * rate -- is left UNDEFINED rather than zeroed.
654 */
656 const bool upper = opt.method != "mw.lower";
657 // Fully qualified: the member `lqn` shadows the namespace of the same
658 // name inside this class, so an unqualified `lqn::` would not compile.
659 const ::line::lqn::LqnBoxBounds<T> b = ::line::lqn::lqn_boxbounds(lqn);
661 const std::size_t N = lqn.nidx;
662 auto blank = [&](std::vector<T>& v, std::vector<bool>& d) {
663 v.assign(N + 1, Tzero());
664 d.assign(N + 1, false);
665 };
666 blank(s.QN, s.defined_Q);
667 blank(s.UN, s.defined_U);
668 blank(s.RN, s.defined_R);
669 blank(s.TN, s.defined_T);
670 blank(s.AN, s.defined_A);
671 blank(s.WN, s.defined_W);
672 for (std::size_t i = 1; i <= N; ++i) {
673 if (b.defined_T[i]) {
674 s.TN[i] = upper ? b.TN_up[i] : b.TN_lo[i];
675 s.defined_T[i] = true;
676 }
677 if (b.defined_U[i]) {
678 s.UN[i] = upper ? b.UN_up[i] : b.UN_lo[i];
679 s.defined_U[i] = true;
680 }
681 }
682 s.iterations = 0;
683 s.converged = true; // a bound needs no fixed point to have converged to
684 s.is_bound = true;
685 return s;
686 }
687
688 std::size_t nlayers() const { return ensemble.size(); }
689 const std::vector<qn::Layer<T>>& layers() const { return ensemble; }
690
691 /**
692 * Port of `LayeredNetwork.layerBlocks`: the block-diagonal layout of the
693 * layers in the aggregate (station x class) view.
694 *
695 * Available as soon as the solver is constructed -- `build_layers` runs in
696 * the constructor -- which is what lets SolverENV compare the shapes of its
697 * stages before solving any of them.
698 */
701 const std::size_t E = ensemble.size();
702 b.roff.assign(E, 0);
703 b.coff.assign(E, 0);
704 b.msz.assign(E, 0);
705 b.ksz.assign(E, 0);
706 for (std::size_t e = 0; e < E; ++e) {
707 b.roff[e] = b.M;
708 b.coff[e] = b.K;
709 b.msz[e] = ensemble[e].nstations;
710 b.ksz[e] = ensemble[e].nclasses;
711 b.M += b.msz[e];
712 b.K += b.ksz[e];
713 }
714 return b;
715 }
716
717 /**
718 * Port of `LayeredNetwork.initFromMarginal`: split an aggregate (M x K)
719 * mean queue-length matrix into per-layer blocks and warm-start each layer
720 * from its own.
721 *
722 * WHY THIS IS RECORDED AND NOT APPLIED HERE, which is the trap the JAR and
723 * python ports both hit: the layered fixed point RESETS its layers as it
724 * converges, so a warm start installed before the solve does not survive
725 * it. The blocks are therefore kept and replayed by `run_layer_transient`,
726 * i.e. after the fixed point and immediately before each layer's transient
727 * -- the steady solve ignores an initial state, the transient does not.
728 *
729 * THE BLOCK IS NOT ROUNDED. The reference rounds a replayed block onto the
730 * integer lattice for every layer engine EXCEPT the fluid one, and the
731 * layered transient exists only over fluid layers (`require_transient_ready`
732 * refuses the rest), so the branch that would round has no reachable case
733 * here. Rounding a fluid state would quantize the very quantity SolverENV's
734 * fixed point is iterating on.
735 *
736 * An empty matrix clears the warm start, so a caller can put the layers
737 * back on their default initial state without rebuilding the solver.
738 */
740 const std::size_t E = ensemble.size();
741 layer_tran_init.assign(E, std::vector<double>());
742 if (n.empty()) return;
743 const LnLayerBlocks b = layer_blocks();
744 if (n.rows() != b.M || n.cols() != b.K)
745 throw InputError(
746 "SolverLN::init_from_marginal: the marginal is " + std::to_string(n.rows()) + "x" +
747 std::to_string(n.cols()) + " where the block-diagonal union of the layers is " +
748 std::to_string(b.M) + "x" + std::to_string(b.K) +
749 "; the aggregate view is layerBlocks', not the LQN element count");
750 for (std::size_t e = 0; e < E; ++e) {
751 const qn::Layer<T>& L = ensemble[e];
753 std::vector<double> y(lay.nstates, 0.0);
754 bool any = false;
755 for (std::size_t i = 0; i < L.nstations && i < b.msz[e]; ++i)
756 for (std::size_t r = 0; r < L.nclasses && r < b.ksz[e]; ++r) {
757 if (!lay.enabled[i][r]) continue;
758 const double q = std::max(0.0, n(b.roff[e] + i, b.coff[e] + r));
759 y[lay.qidx[i][r]] = q;
760 if (q > 0.0) any = true;
761 }
762 // An all-zero block is NOT a warm start: it is the absence of one,
763 // and installing it would empty a closed layer whose population the
764 // fixed point conserves. The reference reaches the same place by
765 // never calling initFromMarginal before the first stage solve.
766 if (any) layer_tran_init[e] = y;
767 }
768 }
769
770 /** Install the explicit output grid of the layered transient; see `LnOptions::tran_grid`. */
771 void set_tran_grid(const std::vector<double>& g) { opt.tran_grid = g; }
772
773 /** Diagnostic access to the per-iteration layer results, for the regression. */
774 const std::vector<std::vector<LayerResult<T>>>& iteration_results() const { return results; }
775 const std::vector<T>& state_servt() const { return servt; }
776 const std::vector<T>& state_residt() const { return residt; }
777 const std::vector<T>& state_tput() const { return tput; }
778 const std::vector<T>& state_thinkt() const { return thinkt; }
779 const std::vector<T>& state_callservt() const { return callservt; }
780 const std::vector<T>& state_callresidt() const { return callresidt; }
781 const std::vector<Distrib<T>>& state_servtproc() const { return servtproc; }
782 const std::vector<Distrib<T>>& state_thinktproc() const { return thinktproc; }
783 const std::vector<Distrib<T>>& state_callservtproc() const { return callservtproc; }
784 const std::vector<T>& state_util() const { return util; }
785 /** The method the layers were built for: "srvn.ph", "srvn.cs" or "moment3". */
786 const std::string& state_lnmethod() const { return lnmethod; }
787 /**
788 * The interlock tracking method this run USES: "ilrate", "refpath" or
789 * "none". Recorded by build_layers and, when no layer merged a reference
790 * chain, demoted from "refpath" to "ilrate" at the end of construct.
791 */
792 const std::string& state_interlock_method() const { return interlockMethod; }
793 /** Per reference-path stage key, the stage mean the last iteration installed. */
794 const std::vector<double>& state_refpath_stages() const { return refpath_stage_prev; }
795
796 /**
797 * Port of SolverLN.interlockMethodFor: the method a solve would use, asked
798 * through the same query build_layers asks. Like the reference it does not
799 * see the construct-time demotion of a refpath that merged nothing; that
800 * recorded answer is state_interlock_method().
801 */
803 return ln_interlock_method(lqn, opt.interlock_method, opt.interlocking, lnmethod);
804 }
805
806private:
807 // state
810
811 std::vector<qn::Layer<T>> ensemble; ///< compacted, layer e is ensemble[e]
812 std::vector<long> idxhash; ///< (nidx+1) element -> layer index, -1 if none
813 std::vector<bool> ignore; ///< (nidx+1) element in a component with no ref task
814
815 // update maps, rows of [layerElement, elementOrCall, node, class]
816 struct UpdRow {
817 std::size_t idx, aidx, node, cls;
818 };
819 std::vector<UpdRow> servt_map, thinkt_map, call_map, actthinkt_map;
820 /**
821 * Source arrival of an async call class, `aidx` holding the call index.
822 *
823 * Kept apart from call_map because the two update different things about
824 * the same class: call_map carries the SERVICE at the server replicas,
825 * this one the arrival RATE at the Source, which is the caller activity's
826 * throughput and so moves with the fixed point. Entry open arrivals have
827 * no row here at all -- theirs is exogenous and never reseeded.
828 */
829 std::vector<UpdRow> arv_call_map;
830 /** [idx, tidx_caller, eidx, nodefrom, nodeto, classfrom, classto] */
831 struct RouteRow {
832 std::size_t idx, tidx_caller, eidx, nodefrom, nodeto, cfrom, cto;
833 };
834 std::vector<RouteRow> route_map;
835 std::vector<std::size_t> unique_route_idx;
836 std::vector<std::size_t> route_reset, svc_reset;
837
838 /** The interlock tracking method in force, resolved once by build_layers. */
839 std::string interlockMethod = "ilrate";
840 /**
841 * Reference-path stage rows, refpath_classes_updmap of the reference:
842 * [idx, node, class] name the stage, `kind` 1 = residt(elem), 2 =
843 * callresidt(elem)/entryvisits(fromentry), 3 = threadWait(elem), each term
844 * multiplied by `sign`.
845 */
846 struct RefPathRow {
847 std::size_t idx, node, cls;
848 int kind;
849 std::size_t elem;
850 double sign;
851 std::size_t fromentry;
852 };
853 std::vector<RefPathRow> refpath_map;
854 /** Unique (idx, node, class) of refpath_map, sorted, and each row's key. */
855 std::vector<std::array<std::size_t, 3>> refpath_keys;
856 std::vector<std::size_t> refpath_group;
857 /** Per key, the previous iteration's stage mean; NaN until the first update. */
858 std::vector<double> refpath_stage_prev;
859 std::vector<T> refpath_stage_prev_v;
860 /**
861 * Per layer ELEMENT, the per-class normalising TASK class (1-based, 0 = use
862 * the chain's reference class), attribute.normclass of the reference. Only
863 * the refpath layers fill it; see ln_layer_refcell.
864 */
865 std::map<std::size_t, std::vector<std::size_t>> layer_normclass;
866 /** Invocations of an entry per invocation of its task; starts at 1. */
867 std::vector<T> entryvisits;
868
869 Matrix<double> njobs; ///< (NT+1 x NT+1) population of caller tidx in layer idx
870
871 // per-element metric state
872 std::vector<T> servt, residt, tput, util, thinkt;
873 std::vector<T> callservt, callresidt;
874 /**
875 * Phase-2 state, the reference's hasPhase2 / servt_ph1 / servt_ph2 /
876 * prOvertake. The split is of `servt` and exists only while `has_phase2`;
877 * `prOvertake` is indexed by entry NUMBER 1..nentries, as the reference
878 * indexes it, not by element index.
879 */
880 bool has_phase2 = false;
881 std::vector<T> servt_ph1, servt_ph2, prOvertake;
882 /** Last entry service and its phase-2 part, resolve_entry_service -> overtaking split. */
883 std::vector<T> entry_servt_last, entry_servt_ph2;
884 /** Tasks whose own layer was collapsed to one representative replica. */
885 std::set<std::size_t> single_replica_tasks;
886 std::vector<Distrib<T>> servtproc, thinkproc, thinktproc, tputproc, callservtproc;
887 /**
888 * `moment3` state: the response-time CDF of each activity (by element) and
889 * of each call, the fitted per-entry law and the CDF read off it. All empty
890 * under the default method, which never forms a distribution at all.
891 */
892 std::vector<fluid::FluidPassage> servtcdf, callservtcdf;
893 std::vector<LnCdf> entrycdfrespt;
894 std::vector<mam::AphPair<T>> entryproc;
895 /** Per layer, the response-time CDFs of the converged pass; a lazy cache. */
896 std::vector<std::vector<std::vector<fluid::FluidPassage>>> cdf_repo;
897 std::vector<double> servt_prev, residt_prev, tput_prev, thinkt_prev;
898 std::vector<double> callservt_prev, callresidt_prev;
899 std::vector<T> servt_prev_v, residt_prev_v, tput_prev_v, thinkt_prev_v, callservt_prev_v;
900 Matrix<T> servtmatrix; ///< (nidx+ncalls) x (nidx+ncalls) entry reachability
901
902 // interlock tables
903 Matrix<T> il_all, il_ph1;
904 std::vector<std::vector<std::size_t>> il_common_entries, il_src_all, il_src_ph2;
905 /**
906 * Per-layer interlock matrix of Franks (1999), Eq. (4.7), CLASS-indexed. A host layer
907 * whose engine carries the correction inside its own MVA takes the matrix instead of
908 * having its residence times scaled after the fact; empty for every other layer.
909 */
910 std::vector<std::vector<std::vector<double>>> layer_interlock;
911 std::vector<double> il_num_sources;
912
913 std::vector<std::vector<LayerResult<T>>> results; ///< [iteration][layer]
914 std::vector<Matrix<T>> layer_init_sol;
915 /**
916 * Per layer, the ODE state `init_from_marginal` recorded, replayed by
917 * `run_layer_transient`. Empty for a layer with no warm start.
918 */
919 std::vector<std::vector<double>> layer_tran_init;
920 /**
921 * Per layer, the fork-join transform SolverMVA solves instead of the layer
922 * itself (inactive for a layer with no fork), and the auxiliary arrival
923 * rates its fixed point iterates on. See fj_mmt.h.
924 *
925 * `fj_lambda` is the `self.fjForkLambda` of the reference and is deliberately
926 * NOT reset between outer iterations: the transform is rebuilt cold on every
927 * outer pass but the iterate it converged to last time is a far better
928 * starting point than FineTol, and the reference warm-starts it the same way
929 * (options.config.fj_warmstart, true by default).
930 */
931 std::vector<mva::FjMmt<T>> fj_tr;
932 std::vector<std::vector<T>> fj_lambda;
933 double relax_omega = 1.0;
934 /** Live only while a NOISY layer engine is in force; null otherwise. */
935 std::shared_ptr<LnStochController<T>> stoch_ctl;
936 long averagingstart = -1;
937 bool hasconverged = false;
938 /** method='moment3': true once the distribution pass has formed the entry laws. */
939 bool moment_pass_done = false;
940 std::vector<double> maxitererr;
941 int iterations_done = 0;
942 bool did_converge = false;
943
944 std::size_t NT() const { return lqn.tshift + lqn.ntasks; }
945
946 static T Tzero() { return num_traits<T>::from_int(0); }
947 static T Tone() { return num_traits<T>::from_int(1); }
948 static double dbl(const T& x) { return num_traits<T>::to_double(x); }
949
950 // -----------------------------------------------------------------------
951 // construct
952 // -----------------------------------------------------------------------
953 void construct() {
954 // Forwarding rewrite, SolverLN.m:162-164: flatten every forwarding
955 // chain into caller-side pseudo rendezvous calls BEFORE anything else
956 // reads lqn, so build_layers/detect_phase2/reject_unsupported all see
957 // plain rendezvous arcs. The raw FWD calls survive in the struct but
958 // must not contribute blocking after this point (lqn_helpers.h).
960 reject_unsupported();
961 detect_phase2();
962
963 const std::size_t N = lqn.nidx;
964 ignore.assign(N + 1, false);
965 // weakly connected components of graph + graph'; a component with no
966 // reference task is unreachable and its elements are ignored
967 std::vector<long> comp(N + 1, -1);
968 long ncomp = 0;
969 for (std::size_t v = 1; v <= N; ++v) {
970 if (comp[v] >= 0) continue;
971 std::vector<std::size_t> stack{v};
972 comp[v] = ncomp;
973 while (!stack.empty()) {
974 const std::size_t u = stack.back();
975 stack.pop_back();
976 for (std::size_t w : lqn.graph.succ(u))
977 if (comp[w] < 0) {
978 comp[w] = ncomp;
979 stack.push_back(w);
980 }
981 for (std::size_t w : lqn.graph.pred(u))
982 if (comp[w] < 0) {
983 comp[w] = ncomp;
984 stack.push_back(w);
985 }
986 }
987 ++ncomp;
988 }
989 if (ncomp > 1) {
990 std::vector<bool> has_ref(ncomp, false);
991 for (std::size_t t = 1; t <= lqn.ntasks; ++t) {
992 const std::size_t tidx = lqn.tshift + t;
993 if (lqn.sched[tidx] == SchedStrategy::REF) has_ref[comp[tidx]] = true;
994 }
995 for (std::size_t v = 1; v <= N; ++v)
996 if (!has_ref[comp[v]]) ignore[v] = true;
997 }
998
999 servtproc.assign(N + 1, Distrib<T>::disabled_dist());
1000 thinkproc.assign(N + 1, Distrib<T>::disabled_dist());
1001 thinktproc.assign(N + 1, Distrib<T>::disabled_dist());
1002 tputproc.assign(N + 1, Distrib<T>::disabled_dist());
1003 for (std::size_t i = 1; i <= N; ++i) {
1004 servtproc[i] = lqn.hostdem[i];
1005 thinkproc[i] = lqn.think[i];
1006 }
1007 callservtproc.assign(lqn.ncalls + 1, Distrib<T>::disabled_dist());
1008 for (std::size_t c = 1; c <= lqn.ncalls; ++c)
1009 callservtproc[c] = lqn.hostdem[lqn.callpair_dst[c]];
1010
1011 // Invocations of an entry per invocation of its task, filled by
1012 // resolve_entry_service; 1 because iteration 1 runs before any measurement
1013 entryvisits.assign(N + 1, Tone());
1014 refpath_map.clear();
1015 layer_normclass.clear();
1016
1017 njobs = Matrix<double>(NT() + 1, NT() + 1, 0.0);
1018 build_layers();
1019
1020 // layers whose routing or service must be re-derived after an update
1021 std::vector<bool> rr(NT() + 1, false), sr(NT() + 1, false);
1022 for (const RouteRow& r : route_map) rr[r.idx] = true;
1023 for (const UpdRow& r : thinkt_map) sr[r.idx] = true;
1024 for (const UpdRow& r : call_map) sr[r.idx] = true;
1025 // refpath REMOVES the think-time row of every merged caller and the
1026 // client-side call row of every re-expanded descent, so a layer whose only
1027 // per-iteration change was one of those would otherwise never be
1028 // refresh_rates()'d and would freeze at its build-time seed. Its stage
1029 // rows put it back.
1030 for (const RefPathRow& r : refpath_map) sr[r.idx] = true;
1031 // a moved arrival rate reweights the Sink closure, which only
1032 // refresh_chains rebuilds, so these layers need the routing reset
1033 for (const UpdRow& r : arv_call_map) rr[r.idx] = true;
1034 for (std::size_t i = 1; i <= NT(); ++i) {
1035 if (rr[i] && idxhash[i] >= 0) route_reset.push_back(std::size_t(idxhash[i]));
1036 if (sr[i] && idxhash[i] >= 0) svc_reset.push_back(std::size_t(idxhash[i]));
1037 }
1038 std::sort(route_reset.begin(), route_reset.end());
1039 route_reset.erase(std::unique(route_reset.begin(), route_reset.end()), route_reset.end());
1040 std::sort(svc_reset.begin(), svc_reset.end());
1041 svc_reset.erase(std::unique(svc_reset.begin(), svc_reset.end()), svc_reset.end());
1042
1043 // The reference-path stage keys, unique((idx, node, class), 'rows').
1044 refpath_keys.clear();
1045 refpath_group.clear();
1046 for (const RefPathRow& r : refpath_map) refpath_keys.push_back({r.idx, r.node, r.cls});
1047 std::sort(refpath_keys.begin(), refpath_keys.end());
1048 refpath_keys.erase(std::unique(refpath_keys.begin(), refpath_keys.end()),
1049 refpath_keys.end());
1050 for (const RefPathRow& r : refpath_map) {
1051 const std::array<std::size_t, 3> key{r.idx, r.node, r.cls};
1052 refpath_group.push_back(std::size_t(
1053 std::lower_bound(refpath_keys.begin(), refpath_keys.end(), key) -
1054 refpath_keys.begin()));
1055 }
1056 if (refpath_map.empty() && interlockMethod == "refpath") {
1057 // Not one layer merged a reference chain, so the build is the one
1058 // 'none' produces, and so would the solve be: post() suppresses the
1059 // rate discount whenever the method is 'refpath'. Switching a
1060 // correction off and declining to replace it is strictly worse than
1061 // either method alone, so the method is demoted here, before init()
1062 // decides whether to build the interlock tables.
1063 interlockMethod = "ilrate";
1064 }
1065 refpath_stage_prev.assign(refpath_keys.size(), std::numeric_limits<double>::quiet_NaN());
1066 refpath_stage_prev_v.assign(refpath_keys.size(), Tzero());
1067 }
1068
1069 /** Reject, by name, every construct this port does not implement. */
1070 // Phase-2 detection, SolverLN.m:165-173. Deliberately NOT inside
1071 // reject_unsupported: that method is const because it only refuses, and
1072 // detection sets state. Marking the flag mutable would have compiled and
1073 // left a validator that silently mutates the solver.
1074 void detect_phase2() {
1075 has_phase2 = false;
1076 for (std::size_t a = 1; a <= lqn.nacts; ++a)
1077 if (lqn.actphase[a] > 1) has_phase2 = true;
1078 }
1079
1080 /**
1081 * True when `eidx` is the destination of a raw forwarding call.
1082 *
1083 * buildLayersRecursive.m:186-192 keeps this test even though
1084 * lqn_fwd_rendezvous has already run, and so must this port:
1085 * the rewrite walks chains out of SYNC calls only, so an entry whose
1086 * forwarder is reached asynchronously gets no pseudo arc and would
1087 * otherwise be dropped from the layer along with all of its work.
1088 */
1089 bool is_fwd_target(std::size_t eidx) const {
1090 for (std::size_t c = 1; c <= lqn.ncalls; ++c)
1091 if (lqn.calltype[c] == CallType::FWD && lqn.callpair_dst[c] == eidx) return true;
1092 return false;
1093 }
1094
1095 /**
1096 * Total exogenous rate into the entries of task `tidx`, zero unless the arrival is
1097 * the ONLY way in.
1098 *
1099 * A task nobody calls has no task layer, so update_think_times never sets its
1100 * surrogate delay and its caller class cycles against an Immediate one. Carrying the
1101 * arrival as an open stream ON TOP of that unthrottled chain loads the host twice:
1102 * lqn_open_arrival read the processor at 0.68 where lqns, lqsim and LDES all give
1103 * 0.32. The chain is the representation that honours the thread pool, so build_layer
1104 * drops the stream for these tasks and update_think_times closes the chain on this
1105 * rate, as the reference does for a forwarding target. With a caller or a forwarding
1106 * source the stream needs a class of its own and this returns 0.
1107 */
1108 double open_arrival_rate_of(std::size_t tidx) const {
1109 if (lqn.isref[tidx]) return 0.0;
1110 for (std::size_t e : lqn.entriesof[tidx])
1111 if (lqn.issynccaller.any_col(e) || lqn.isasynccaller.any_col(e) || is_fwd_target(e))
1112 return 0.0;
1113 double rate = 0.0;
1114 for (std::size_t e : lqn.entriesof[tidx]) {
1115 if (!lqn.has_arrival[e]) continue;
1116 const double m = dbl(lqn.arrival[e].mean);
1117 if (std::isfinite(m) && m > GlobalConstants::FineTol) rate += 1.0 / m;
1118 }
1119 return rate;
1120 }
1121
1122 /** True when layer `e` carries a Cache node. */
1123 bool has_cache_node(std::size_t e) const {
1124 for (std::size_t n = 0; n < ensemble[e].nodes.size(); ++n)
1125 if (ensemble[e].nodes[n].nodetype == NodeType::Cache) return true;
1126 return false;
1127 }
1128
1129 /** True when `aidx` is an activity bound to an entry nobody calls synchronously. */
1130 bool async_only_activity(std::size_t aidx) const {
1131 if (aidx <= lqn.ashift || aidx > lqn.ashift + lqn.nacts) return false;
1132 for (std::size_t e = 1; e <= lqn.nentries; ++e) {
1133 const std::size_t eidx = lqn.eshift + e;
1134 if (lqn.graph.get(eidx, aidx) == Tzero()) continue;
1135 return lqn.isasynccaller.any_col(eidx) && !lqn.issynccaller.any_col(eidx);
1136 }
1137 return false;
1138 }
1139
1140 /**
1141 * Reject, by name, every construct this port does not implement.
1142 *
1143 * NOTHING IS LEFT. Forwarding, asynchronous calls, entry open arrivals,
1144 * admission constraints, cache tasks and setup tasks are each solved.
1145 * The method is kept rather than deleted because it is the place a new
1146 * refusal belongs, and because a layer-build refusal must be raised HERE,
1147 * before construct() has read anything out of the struct, rather than
1148 * halfway through building a layer.
1149 */
1150 void reject_unsupported() const {}
1151
1152 /**
1153 * The interlock tracking method, resolved ONCE for the encoding just chosen,
1154 * for the same reason lnmethod is: what update_populations and init_interlock
1155 * do can never disagree with what was built into the layers.
1156 */
1157 void resolve_interlock_method() {
1158 const LnInterlockChoice ch =
1159 ln_interlock_method(lqn, opt.interlock_method, opt.interlocking, lnmethod);
1160 interlockMethod = ch.method;
1161 if (!ch.why.empty()) std::cerr << "[LINE] Warning: " << ch.why << "\n";
1162 }
1163
1164 // -----------------------------------------------------------------------
1165 // buildLayers
1166 // -----------------------------------------------------------------------
1167 void build_layers() {
1168 // Method resolution. A method name carries both the LAYERING and the
1169 // ENCODING: "srvn.ph" replaces the routing encoding of the activity graph
1170 // by a composed phase-type server law, "srvn" is the alias that takes it
1171 // where it can serve the model and "srvn.cs" otherwise. The choice is
1172 // made ONCE, here, and every later dispatch reads lnmethod.
1173 // See _kb/06-solver-catalog.md (LN section).
1174 const std::string requested = ln_requested_method(opt.method);
1175 assert_call_groups(requested == "flat.cs");
1176 if (requested == "flat.cs") {
1177 lnmethod = "flat.cs";
1178 resolve_interlock_method();
1179 build_flat_layer();
1180 return;
1181 }
1182 if (requested == "flat.ph") {
1183 // the squashed layering with the composed law: ONE submodel holding
1184 // every server, and a caller visiting each of them once per
1185 // invocation. The feature gate is the srvn.ph one plus the refusals a
1186 // single submodel carries -- see ph_flat_server_set.
1187 ph_laws_ready = false;
1188 lnmethod = "flat.ph";
1189 build_layers_ph(true);
1190 return;
1191 }
1192 if (requested == "srvn.ph" || requested == "srvn") {
1193 ph_laws_ready = false;
1194 if (requested == "srvn.ph" || probe_srvn_ph()) {
1195 lnmethod = "srvn.ph";
1196 build_layers_ph();
1197 return;
1198 }
1199 }
1200 lnmethod = (requested == "moment3") ? "moment3" : "srvn.cs";
1201 resolve_interlock_method();
1202 std::vector<qn::Layer<T>> raw(NT() + 1);
1203 std::vector<bool> present(NT() + 1, false);
1204
1205 for (std::size_t hidx = 1; hidx <= lqn.nhosts; ++hidx) {
1206 if (ignore[hidx]) continue;
1207 build_layer(raw[hidx], {hidx}, lqn.tasksof[hidx], true, false);
1208 present[hidx] = true;
1209 }
1210 for (std::size_t t = 1; t <= lqn.ntasks; ++t) {
1211 const std::size_t tidx = lqn.tshift + t;
1212 if (ignore[tidx] || lqn.isref[tidx]) continue;
1213 bool any_caller = lqn.iscaller.any_row(tidx) || lqn.iscaller.any_col(tidx);
1214 if (!any_caller) continue;
1215 // tasks that call some entry of tidx
1216 std::vector<std::size_t> callers;
1217 for (std::size_t ct = 1; ct <= lqn.ntasks; ++ct) {
1218 const std::size_t c = lqn.tshift + ct;
1219 bool calls = false;
1220 for (std::size_t e : lqn.entriesof[tidx])
1221 if (lqn.iscaller.get(c, e)) calls = true;
1222 if (calls) callers.push_back(c);
1223 }
1224 if (callers.empty()) continue;
1225 build_layer(raw[tidx], {tidx}, callers, false, false);
1226 present[tidx] = true;
1227 }
1228
1229 idxhash.assign(lqn.nidx + 1, -1);
1230 long next = 0;
1231 for (std::size_t i = 1; i <= NT(); ++i)
1232 if (present[i]) {
1233 idxhash[i] = next++;
1234 ensemble.push_back(std::move(raw[i]));
1235 }
1236 layer_init_sol.assign(ensemble.size(), Matrix<T>());
1237 build_fork_views();
1238 }
1239
1240 /**
1241 * The processors and called tasks that become stations of the flat layer.
1242 *
1243 * Squashing is refused rather than approximated where an element carries
1244 * state that only a submodel of its own can hold: a REPLICATED element
1245 * would need one station per copy inside a layer whose routing addresses it
1246 * once, a CACHE task needs the Cache node in the host layer its reads queue
1247 * at, and a SETUP task's delay-off belongs to the station that powers down.
1248 */
1249 /**
1250 * Reject a routed call group under any layering or layer solver that cannot
1251 * carry it.
1252 *
1253 * Two conditions, and both are refusals rather than degradations. The
1254 * squashed layering is needed because under `srvn` each target lives in a
1255 * submodel of its own and is replaced, in the caller's submodel, by a
1256 * surrogate delay -- no node ever has arcs to more than one of them, so
1257 * there is nothing to dispatch among. A layer solver that resolves the
1258 * strategy from the STATE is needed because `refresh_routing` expands
1259 * RROBIN and JSQ into a uniform probability split for the matrix solvers,
1260 * and returning that split under a round-robin label misreports a
1261 * deterministic policy as a coin. In this port only `ssa` resolves them.
1262 */
1263 void assert_call_groups(bool flat) const {
1264 if (lqn.callgroups.empty()) return;
1265 if (!flat)
1266 throw UnsupportedError(
1267 "Call groups routed by a routing strategy require the squashed layering; use "
1268 "method='flat'. Under srvn the targets never share a submodel, so the dispatch "
1269 "order cannot be represented.");
1270 if (opt.layer_solver != "ssa")
1271 throw UnsupportedError(
1272 "Routed call groups need a layer solver that resolves the strategy from the "
1273 "state; set layer_solver='ssa'. MVA, NC and FLD read the routing matrix, into "
1274 "which refresh_routing has expanded the strategy as a uniform split, and would "
1275 "return that split under a round-robin or JSQ label.");
1276 }
1277
1278 std::vector<std::size_t> flat_server_set() const {
1279 std::vector<std::size_t> servers;
1280 for (std::size_t i = 1; i <= NT(); ++i) {
1281 if (lqn.repl[i] > 1.0)
1282 throw UnsupportedError(
1283 "Flat layering does not support replicated processors or tasks, use the "
1284 "default 'srvn' layering.");
1285 if (lqn.iscache[i])
1286 throw UnsupportedError(
1287 "Flat layering does not support cache tasks, use the default 'srvn' "
1288 "layering.");
1289 if (lqn.hassetup[i])
1290 throw UnsupportedError(
1291 "Flat layering does not support setup tasks, use the default 'srvn' "
1292 "layering.");
1293 }
1294 for (std::size_t hidx = 1; hidx <= lqn.nhosts; ++hidx)
1295 if (!ignore[hidx] && !lqn.tasksof[hidx].empty()) servers.push_back(hidx);
1296 for (std::size_t t = 1; t <= lqn.ntasks; ++t) {
1297 const std::size_t tidx = lqn.tshift + t;
1298 if (ignore[tidx] || lqn.isref[tidx]) continue;
1299 if (!lqn.iscaller.any_row(tidx) && !lqn.iscaller.any_col(tidx)) continue;
1300 bool has_task_caller = false;
1301 for (std::size_t eidx : lqn.entriesof[tidx])
1302 for (std::size_t c = 1; c <= lqn.ntasks; ++c)
1303 if (lqn.iscaller.get(lqn.tshift + c, eidx)) has_task_caller = true;
1304 if (has_task_caller) servers.push_back(tidx);
1305 }
1306 if (servers.empty())
1307 throw InputError(
1308 "Flat layering found no server: the model has no processor with tasks.");
1309 return servers;
1310 }
1311
1312 /**
1313 * Build the single squashed layer holding every processor and called task.
1314 *
1315 * Every served element resolves to layer 0, which is at once the host layer
1316 * and the task layer, so a consumer that walks `idxhash` finds the same
1317 * network for a processor and for a task and must read the station off
1318 * `server_idx_of` rather than off `serverIdx`.
1319 */
1320 void build_flat_layer() {
1321 const std::vector<std::size_t> flat_servers = flat_server_set();
1322 std::vector<std::size_t> flat_callers;
1323 for (std::size_t t = 1; t <= lqn.ntasks; ++t) {
1324 const std::size_t tidx = lqn.tshift + t;
1325 if (!ignore[tidx]) flat_callers.push_back(tidx);
1326 }
1327 qn::Layer<T> layer;
1328 build_layer(layer, flat_servers, flat_callers, false, true);
1329 ensemble.clear();
1330 ensemble.push_back(std::move(layer));
1331 idxhash.assign(lqn.nidx + 1, -1);
1332 for (std::size_t s : flat_servers) idxhash[s] = 0;
1333 layer_init_sol.assign(ensemble.size(), Matrix<T>());
1334 build_fork_views();
1335 }
1336
1337 /**
1338 * Build, for every layer that has a fork, the transformed model SolverMVA
1339 * actually solves, and seed its auxiliary arrival rates.
1340 */
1341 void build_fork_views() {
1342 fj_tr.assign(ensemble.size(), mva::FjMmt<T>());
1343 fj_lambda.assign(ensemble.size(), {});
1344 for (std::size_t e = 0; e < ensemble.size(); ++e) {
1345 if (!ensemble[e].has_fork()) continue;
1346 // fj_mmt mints its own Source/Sink pair and would silently overwrite
1347 // sourceIdx/sinkNode, detaching the open stream already routed there.
1348 // NOT a parity gap: the reference does not solve this combination
1349 // either, it fails inside the layer solve with "Arrays have
1350 // incompatible sizes" (checked 2026-07-29 on a fork+async model).
1351 // Refusing by name is the better answer, so this stays.
1352 if (ensemble[e].sourceIdx != 0)
1353 throw UnsupportedError(
1354 "SolverLN: layer '" + ensemble[e].name +
1355 "' carries both an AND fork and an open stream (an async call or an entry "
1356 "arrival); the fork-join transform needs a Source of its own");
1357 fj_tr[e] = mva::fj_mmt(ensemble[e]);
1358 fj_lambda[e].assign(fj_tr[e].V.classes.size() + 1,
1359 num_traits<T>::from_double(GlobalConstants::FineTol));
1360 }
1361 }
1362
1363 /**
1364 * Port of matlab/src/lang/layered/lqn_dep_layer_handle.m: lift a
1365 * per-operand rate handle onto the classes of a layer station.
1366 *
1367 * COLS[j] lists the layer classes (0-based) through which operand j occupies
1368 * the station; CHAINCOLS is the same list in the CHAIN index space, filled
1369 * once refresh_chains has run. Solvers evaluate the handle in either space
1370 * -- CTMC and the exact recursions pass a per-class vector, the AMVA and NC
1371 * chain recursions a per-chain one -- so the handle reads the LENGTH of what
1372 * it is given to decide, aggregates the operand populations in that space,
1373 * and answers a vector of the same length, since the caller indexes the
1374 * answer with the index it passed in. An index belonging to no operand keeps
1375 * the neutral scaling 1.
1376 */
1377 static lang::CdScaling<T> layer_dep_handle(
1378 const lang::CdScaling<T>& f, const std::vector<std::vector<std::size_t>>& cols,
1379 const std::shared_ptr<std::vector<std::vector<std::size_t>>>& chaincols, std::size_t R) {
1380 return [f, cols, chaincols, R](const std::vector<T>& n) -> std::vector<T> {
1381 const std::size_t L = n.size();
1382 const std::vector<std::vector<std::size_t>>& use =
1383 (L == R || chaincols->empty()) ? cols : *chaincols;
1384 std::vector<T> nop(use.size(), num_traits<T>::from_int(0));
1385 for (std::size_t j = 0; j < use.size(); ++j)
1386 for (std::size_t k = 0; k < use[j].size(); ++k)
1387 if (use[j][k] < L) nop[j] = T(nop[j] + n[use[j][k]]);
1388 const std::vector<T> w = f(nop);
1389 std::vector<T> v(L, num_traits<T>::from_int(1));
1390 if (w.empty()) return v;
1391 for (std::size_t j = 0; j < use.size(); ++j) {
1392 const T wj = w[j < w.size() ? j : w.size() - 1];
1393 for (std::size_t k = 0; k < use[j].size(); ++k)
1394 if (use[j][k] < L) v[use[j][k]] = wj;
1395 }
1396 return v;
1397 };
1398 }
1399
1400 /** Spread a per-operand peak rate onto the layer classes. Twin of layerPeak. */
1401 static std::vector<T> layer_peak(const std::vector<T>& peakPerOperand,
1402 const std::vector<std::vector<std::size_t>>& cols,
1403 std::size_t R) {
1404 std::vector<T> peak(R, num_traits<T>::from_int(1));
1405 if (peakPerOperand.empty()) return peak;
1406 for (std::size_t j = 0; j < cols.size(); ++j) {
1407 const T pj = peakPerOperand[j < peakPerOperand.size() ? j : peakPerOperand.size() - 1];
1408 for (std::size_t k = 0; k < cols[j].size(); ++k)
1409 if (cols[j][k] < R) peak[cols[j][k]] = pj;
1410 }
1411 return peak;
1412 }
1413
1414 /**
1415 * Give every class of a host layer the priority of the task that owns it.
1416 *
1417 * Applies to each processor in idxSet that schedules by priority (HOL, 'pri', ...).
1418 * lqns serves a LARGER task priority first while a class priority of 0 is the highest,
1419 * so the class priority is max(prio on the host) - prio(task). A layer solver with a
1420 * priority arm (MVA, CTMC, ...) then applies it. Without this every class tied at 0
1421 * and a priority processor ran as FCFS. Port of MATLAB ln_host_task_priorities.m.
1422 */
1423 void apply_host_task_priorities(qn::Layer<T>& m, const std::vector<std::size_t>& idxSet) const {
1424 using qn::SchedStrategy;
1425 auto is_prio = [](SchedStrategy s) {
1426 return s == SchedStrategy::HOL || s == SchedStrategy::FCFSPRPRIO ||
1427 s == SchedStrategy::FCFSPIPRIO || s == SchedStrategy::LCFSPRIO ||
1428 s == SchedStrategy::LCFSPRPRIO || s == SchedStrategy::LCFSPIPRIO ||
1429 s == SchedStrategy::PSPRIO || s == SchedStrategy::DPSPRIO ||
1430 s == SchedStrategy::GPSPRIO;
1431 };
1432 const std::size_t R = m.classes.size();
1433 std::vector<std::size_t> owner(R + 1, 0); // 1-based class -> owning task, 0 = none
1434 auto own = [&](std::size_t k, std::size_t t) {
1435 if (k >= 1 && k <= R && owner[k] == 0) owner[k] = t;
1436 };
1437 for (const auto& a : m.attr_tasks) own(a.first, a.second);
1438 for (const auto& a : m.attr_entries) own(a.first, lqn.parent[a.second]);
1439 for (const auto& a : m.attr_activities) own(a.first, lqn.parent[a.second]);
1440 for (const auto& a : m.attr_calls) own(a[0], lqn.parent[a[2]]);
1441 for (std::size_t s : idxSet) {
1442 if (s < 1 || s > lqn.nhosts || !is_prio(lqn.sched[s])) continue;
1443 int pmax = 0;
1444 for (std::size_t t = lqn.tshift + 1; t <= lqn.tshift + lqn.ntasks; ++t)
1445 if (lqn.parent[t] == s) pmax = std::max(pmax, lqn.prio[t]);
1446 for (std::size_t k = 1; k <= R; ++k)
1447 if (owner[k] != 0 && lqn.parent[owner[k]] == s)
1448 m.classes[k - 1].prio = pmax - lqn.prio[owner[k]];
1449 }
1450 }
1451
1452 /**
1453 * Port of buildLayersRecursive.
1454 *
1455 * The layer holds a client Delay carrying the callers' think times and the
1456 * time they spend elsewhere, and `nreplicas` copies of the server station.
1457 * A class is created for every task, entry, activity and call that the
1458 * callers' activity graphs reach, and the graph traversal lays down the
1459 * routing that moves a job between those classes.
1460 */
1461 void build_layer(qn::Layer<T>& m, const std::vector<std::size_t>& idxSet,
1462 const std::vector<std::size_t>& callers, bool ishostlayer, bool flat) {
1463 const T one = Tone();
1464 // The layer key: the model name, the ensemble slot and the column every
1465 // update map is written under. Under `srvn` it is the layer's only
1466 // served element; under `flat` it is the first of them.
1467 const std::size_t idx = idxSet[0];
1468 m.name = flat ? lqn.hashnames[idx] + ".Flat" : lqn.hashnames[idx];
1469 m.flat = flat;
1470
1471 const double rawrepl = lqn.repl[idx];
1472 std::size_t nreplicas = 1;
1473 if (!flat && rawrepl > 1.0 && !callers.empty()) {
1474 // A replicated server is pooled into ONE station instead of one per
1475 // replica when every caller already addresses all of its replicas:
1476 // on a host layer that is the callers replicating in step with it,
1477 // on a task layer it is a declared fan-out covering every replica.
1478 // With no fan-out declared the lookup is 0 < rawrepl and the layer
1479 // materialises the replicas, which is the reference's answer too.
1480 bool reduce = true;
1481 if (ishostlayer) {
1482 for (std::size_t c : callers)
1483 if (lqn.repl[c] != rawrepl) reduce = false;
1484 } else {
1485 for (std::size_t c : callers)
1486 if (lqn.fanout_at(c, idx) < rawrepl) reduce = false;
1487 }
1488 nreplicas = reduce ? 1 : static_cast<std::size_t>(std::llround(rawrepl));
1489 if (reduce && !ishostlayer) single_replica_tasks.insert(idx);
1490 }
1491 const bool reduce_fanout = (nreplicas == 1 && rawrepl > 1.0 && !callers.empty());
1492 const std::vector<double>& mult = lqn.maxmult;
1493
1494 // A client Delay exists unless the layer is a task layer reached only
1495 // by asynchronous callers, which this port refuses upstream.
1496 m.clientIdx = m.add_station(qn::Station<T>{"Clients", NodeType::Delay, SchedStrategy::INF,
1497 std::numeric_limits<double>::infinity(), false, 0});
1498 const std::size_t clientNode = m.node_of_station(m.clientIdx);
1499 m.serverIdx = m.clientIdx + 1;
1500 // One station (times its replicas) per SERVED ELEMENT of this layer.
1501 // Under `srvn` that is one element and `srv[idx] == server`; under
1502 // `flat` every processor and called task of the model is here.
1503 std::vector<std::vector<std::size_t>> srv(lqn.nidx + 1);
1504 std::vector<std::vector<std::size_t>> srvnode(lqn.nidx + 1);
1505 m.server_idx_of.assign(lqn.nidx + 1, 0);
1506 for (std::size_t sidx : idxSet) {
1507 const bool sishost = sidx <= lqn.nhosts;
1508 srv[sidx].resize(nreplicas);
1509 for (std::size_t r = 0; r < nreplicas; ++r) {
1510 qn::Station<T> st;
1511 st.name = r == 0 ? lqn.hashnames[sidx]
1512 : lqn.hashnames[sidx] + "." + std::to_string(r + 1);
1513 // inf-scheduled Queue as Delay node rationale: see _kb/06-solver-catalog.md (cpp port notes: solver_ln.h)
1514 st.nodetype =
1515 lqn.sched[sidx] == SchedStrategy::INF ? NodeType::Delay : NodeType::Queue;
1516 st.sched = lqn.sched[sidx];
1517 // setNumberOfServers no-op rationale: see _kb/06-solver-catalog.md (cpp port notes: solver_ln.h)
1518 st.nservers = lqn.sched[sidx] == SchedStrategy::INF
1519 ? std::numeric_limits<double>::infinity()
1520 : mult[sidx];
1521 st.attr_ishost = flat ? sishost : ishostlayer;
1522 st.attr_idx = sidx;
1523 srv[sidx][r] = m.add_station(st);
1524 srvnode[sidx].push_back(m.node_of_station(srv[sidx][r]));
1525 }
1526 m.server_idx_of[sidx] = srv[sidx][0];
1527 if (sishost)
1528 m.host_stations.push_back(srv[sidx][0]);
1529 else
1530 m.task_stations.push_back(srv[sidx][0]);
1531 }
1532 // the layer's own server, the sole server under `srvn`
1533 const std::vector<std::size_t>& server = srv[idx];
1534 const std::vector<std::size_t>& serverNode = srvnode[idx];
1535
1536 /** Stations of ELEM when it is served in this layer, empty otherwise. */
1537 auto servers_for = [&](std::size_t elem) -> const std::vector<std::size_t>& {
1538 static const std::vector<std::size_t> none;
1539 if (elem >= 1 && elem <= lqn.nidx && !srv[elem].empty()) return srv[elem];
1540 return none;
1541 };
1542 auto server_nodes_for = [&](std::size_t elem) -> const std::vector<std::size_t>& {
1543 static const std::vector<std::size_t> none;
1544 if (elem >= 1 && elem <= lqn.nidx && !srvnode[elem].empty()) return srvnode[elem];
1545 return none;
1546 };
1547 /** True when the processor of task TIDX is a server of this layer. */
1548 auto host_is_server = [&](std::size_t tidx_) {
1549 return !servers_for(lqn.parent[tidx_]).empty();
1550 };
1551 /** True when TIDX calls an entry served by this layer. */
1552 auto is_layer_client = [&](std::size_t tidx_) {
1553 for (std::size_t s : idxSet)
1554 for (std::size_t e : lqn.entriesof[s])
1555 if (lqn.issynccaller.get(tidx_, e)) return true;
1556 return false;
1557 };
1558 /** Declare a class at every server station of this layer. */
1559 auto set_all_servers = [&](std::size_t cl, const Distrib<T>& d) {
1560 for (std::size_t s : idxSet)
1561 for (std::size_t st : srv[s]) m.set_service(st, cl, d);
1562 };
1563
1564 // Routed call groups, resolved from target ENTRIES to the call indices
1565 // that reach them: a group is ONE dispatch with n destinations, so its
1566 // members share a call class and a dispatch class that holds the job at
1567 // the client while the target is picked (callGroupsByCidx in
1568 // buildLayersRecursive.m).
1569 std::vector<std::size_t> group_of_call(lqn.ncalls + 1, 0);
1570 std::vector<std::vector<std::size_t>> group_members(lqn.callgroups.size() + 1);
1571 for (std::size_t g = 0; g < lqn.callgroups.size(); ++g) {
1572 const LqnCallGroup& grp = lqn.callgroups[g];
1573 for (std::size_t tgt : grp.targets)
1574 for (std::size_t cidx : lqn.callsof[grp.caller])
1575 if (lqn.callpair_dst[cidx] == tgt && lqn.calltype[cidx] == CallType::SYNC &&
1576 group_of_call[cidx] == 0) {
1577 group_of_call[cidx] = g + 1;
1578 group_members[g + 1].push_back(cidx);
1579 break;
1580 }
1581 }
1582 // per group: the Router node, the dispatch class and the group class
1583 std::vector<std::size_t> grp_router(lqn.callgroups.size() + 1, 0);
1584 std::vector<std::size_t> grp_dispatch(lqn.callgroups.size() + 1, 0);
1585 std::vector<std::size_t> grp_class(lqn.callgroups.size() + 1, 0);
1586 // (router node, dispatch class, group) triples whose strategy is
1587 // installed once the routing is written
1588 std::vector<std::array<std::size_t, 3>> routed_group_sites;
1589
1590 // Fork/Router/Join construction rationale: see _kb/06-solver-catalog.md (cpp port notes: solver_ln.h)
1591 std::vector<std::size_t> acts_in_caller;
1592 for (std::size_t c : callers)
1593 for (std::size_t a : lqn.actsof[c]) acts_in_caller.push_back(a);
1594 bool hasfork = false, hasjoin = false;
1595 std::size_t maxfanout = 1;
1596 for (std::size_t a : acts_in_caller) {
1597 if (lqn.actposttype[a] == PrecedenceType::POST_AND) hasfork = true;
1598 if (lqn.actpretype[a] == PrecedenceType::PRE_AND) hasjoin = true;
1599 std::size_t nand = 0;
1600 for (std::size_t sx : lqn.graph.succ(a))
1601 if (lqn.actposttype[sx] == PrecedenceType::POST_AND) ++nand;
1602 if (nand > maxfanout) maxfanout = nand;
1603 }
1604 std::size_t forkNode = 0, joinNode = 0, joinStation = 0;
1605 std::vector<std::size_t> forkRouter;
1606 if (hasfork) {
1607 forkNode = m.add_node("Fork_PostAnd", NodeType::Fork, false);
1608 for (std::size_t f = 1; f <= maxfanout; ++f)
1609 forkRouter.push_back(
1610 m.add_node("Fork_PostAnd_" + std::to_string(f), NodeType::Router, true));
1611 }
1612 if (hasjoin) {
1613 qn::Station<T> js;
1614 js.name = "Join_PreAnd";
1615 js.nodetype = NodeType::Join;
1616 js.sched = SchedStrategy::INF;
1617 js.nservers = std::numeric_limits<double>::infinity();
1618 js.attr_ishost = false;
1619 js.attr_idx = 0;
1620 joinStation = m.add_station(js);
1621 joinNode = m.node_of_station(joinStation);
1622 if (forkNode) m.fj.emplace_back(forkNode, joinNode);
1623 }
1624 // A CACHE LAYER is a HOST layer all of whose callers are cache tasks,
1625 // buildLayersRecursive.m:88. The Cache node lives here and not in the
1626 // cache task's own layer: the item lookup is what the task DOES on its
1627 // processor, so the hit and miss branches have to queue at the same
1628 // server the read arrived at.
1629 bool iscachelayer = !flat && ishostlayer && !callers.empty();
1630 for (std::size_t c : callers)
1631 if (!lqn.iscache[c]) iscachelayer = false;
1632 // A SETUP NO LONGER CHANGES HOW A LAYER IS BUILT (buildLayersRecursive.m,
1633 // 2026-08-11). Wiring the setup/delay-off onto the server station routed
1634 // the layer through the open M/G/1-with-setup QBD, which reads the idle
1635 // period off the Poisson rate 1/X and so powers the thread down far more
1636 // often than a CLOSED layer does -- 12.67% below LDES on lqn_setup -- and
1637 // it charged the restart to the ACTIVITY, where it is not host demand.
1638 // The cold start is charged to the ENTRY instead, with the probability
1639 // that the thread was really found down: see setup_charge().
1640 std::size_t cacheNode = 0;
1641 qn::CacheParam<T> cachepar;
1642 if (iscachelayer) {
1643 const std::size_t ct = callers[0];
1644 cachepar.nitems = lqn.nitems[ct];
1645 cachepar.itemcap = lqn.itemcap[ct];
1646 cachepar.replacestrat = lqn.replacestrat[ct];
1647 cacheNode = m.add_node(lqn.hashnames[ct], NodeType::Cache, true);
1648 }
1649
1650 // The open streams this layer carries, decided BEFORE the Source is
1651 // added: m.serverIdx is the client index plus one and replica r is
1652 // m.serverIdx + r, so a station inserted between them would silently
1653 // renumber the servers. Source and Sink therefore come last.
1654 // An async call is declared only in the layer whose server owns the
1655 // called entry, which is serversFor(parent(dst)) in the reference
1656 // (buildLayersRecursive.m:327) and the same test route_sync_call uses.
1657 std::vector<std::size_t> async_here;
1658 for (std::size_t c : callers)
1659 for (std::size_t aidx : lqn.actsof[c])
1660 for (std::size_t cidx : lqn.callsof[aidx])
1661 if (lqn.calltype[cidx] == CallType::ASYNC &&
1662 !servers_for(lqn.parent[lqn.callpair_dst[cidx]]).empty())
1663 async_here.push_back(cidx);
1664 std::vector<std::size_t> open_entries;
1665 for (std::size_t c : callers) {
1666 // An arrival that is the only way into the task is carried by the caller
1667 // CHAIN, not by a stream: that task has no task layer, so nothing ever sets
1668 // its surrogate delay, and a chain cycling against an Immediate one plus a
1669 // stream loads the host twice -- lqn_open_arrival's processor read 0.68
1670 // against 0.32 from lqns, lqsim and LDES. update_think_times closes the
1671 // chain on the known rate instead, as it does for a forwarding target.
1672 if (open_arrival_rate_of(c) > GlobalConstants::FineTol) continue;
1673 for (std::size_t eidx : lqn.entriesof[c])
1674 if (lqn.has_arrival[eidx]) open_entries.push_back(eidx);
1675 }
1676
1677 std::size_t sourceStation = 0, sourceNode = 0, sinkNode = 0;
1678 if (!async_here.empty() || !open_entries.empty()) {
1679 qn::Station<T> src;
1680 src.name = "Source";
1681 src.nodetype = NodeType::Source;
1682 // EXT keeps solver_mva out of both infSET and qSET: the Source is lambda, not a queue
1683 src.sched = SchedStrategy::EXT;
1684 src.nservers = 1.0;
1685 src.attr_ishost = false;
1686 src.attr_idx = 0;
1687 sourceStation = m.add_station(src);
1688 sourceNode = m.node_of_station(sourceStation);
1689 sinkNode = m.add_node("Sink", NodeType::Sink, false);
1690 m.sourceIdx = sourceStation;
1691 m.sinkNode = sinkNode;
1692 }
1693
1694 /** The stack of entry classes at the forks currently open. */
1695 std::vector<std::size_t> forkClassStack;
1696
1697 // class index of each LQN element and each call, 0 when absent
1698 std::vector<std::size_t> cls(lqn.nidx + 1, 0);
1699 std::vector<std::size_t> callcls(lqn.ncalls + 1, 0);
1700 // `<call>.Aux`, present only for a call whose mean count differs from the
1701 // replica count; 0 elsewhere. See the routing in route_sync_call.
1702 std::vector<std::size_t> auxcallcls(lqn.ncalls + 1, 0);
1703
1704 // A caller carries a closed class when its own PROCESSOR is served here
1705 // and something drives it (it is a reference task, is called, is a
1706 // forwarding target or takes an arrival), or when it calls an entry
1707 // served here. Under `srvn` the first disjunct is exactly the host
1708 // layer and the second exactly the task layer, so this reduces to the
1709 // two-branch test it replaces (buildLayersRecursive.m:217).
1710 auto caller_needs_class = [&](std::size_t tidx_caller) {
1711 if (host_is_server(tidx_caller)) {
1712 if (lqn.isref[tidx_caller]) return true;
1713 for (std::size_t e : lqn.entriesof[tidx_caller])
1714 if (lqn.issynccaller.any_col(e) || lqn.isasynccaller.any_col(e) ||
1715 is_fwd_target(e) || lqn.has_arrival[e])
1716 return true;
1717 }
1718 return is_layer_client(tidx_caller);
1719 };
1720 auto caller_acts_visible = [&](std::size_t tidx_caller) {
1721 return host_is_server(tidx_caller) || is_layer_client(tidx_caller);
1722 };
1723 /**
1724 * The population caller TIDX_ contributes to this layer WITHOUT a
1725 * refpath merge, including the infinite-server fallbacks. Shared by the
1726 * first pass, which also records it in njobs, and by resolve_ref_path,
1727 * which compares it against the chain population before any class exists.
1728 */
1729 auto caller_pool = [&](std::size_t tidx_) -> double {
1730 if (njobs(tidx_, idx) != 0.0) return njobs(tidx_, idx);
1731 // representative-replica scaling rationale: see _kb/06-solver-catalog.md (cpp port notes: solver_ln.h)
1732 const bool single_replica = reduce_fanout || single_replica_tasks.count(tidx_) > 0;
1733 double nj = single_replica ? mult[tidx_] : mult[tidx_] * lqn.repl[tidx_];
1734 if (std::isinf(nj)) {
1735 double s = 0.0;
1736 for (std::size_t c = 1; c <= NT(); ++c)
1737 if (lqn.taskgraph.get(c, tidx_) != Tzero()) s += mult[c];
1738 nj = s;
1739 if (std::isinf(nj)) {
1740 double s2 = 0.0;
1741 for (std::size_t c = 1; c <= NT(); ++c)
1742 if (std::isfinite(mult[c])) s2 += mult[c] * lqn.repl[c];
1743 nj = std::min(s2, 1000.0);
1744 }
1745 }
1746 return nj;
1747 };
1748
1749 // ---- reference-path interlocking (interlock_method = 'refpath') -------
1750 // This layer's client is then not one population per caller but ONE
1751 // chain running from the reference task down through every caller it
1752 // reaches, so two callers that are the same REF customers arriving by
1753 // two routes stop being counted as two. The tables say which callers are
1754 // merged (rp_spliced), which sync calls stop being a surrogate delay and
1755 // become an explicit descent (rp_expand), and which entries above or
1756 // between the callers become hop stages on the client Delay (rp_hop).
1757 // See lqn_ref_routes and _kb/06-solver-catalog.md (LN section).
1758 bool rp_on = false;
1759 std::vector<bool> rp_member(lqn.nidx + 1, false); // task: a caller merged into a chain
1760 std::vector<bool> rp_spliced(lqn.nidx + 1, false); // task: merged AND not the chain head
1761 std::vector<bool> rp_hop(lqn.nidx + 1, false); // entry: a stage of the path
1762 std::vector<bool> rp_expand(lqn.ncalls + 1, false); // call: re-expanded, not a delay
1763 std::vector<std::size_t> rp_child_task(lqn.ncalls + 1, 0); // callee task, when a member
1764 std::vector<std::size_t> rp_child_entry(lqn.ncalls + 1, 0); // callee entry, when a hop
1765 std::vector<T> rp_ret_share(lqn.ncalls + 1, Tzero()); // share of the callee's return
1766 std::vector<T> rp_call_mean(lqn.ncalls + 1, Tzero()); // invocations per execution
1767 std::vector<std::vector<std::size_t>> rp_hop_children(lqn.nidx + 1); // descent order
1768 std::vector<double> rp_hop_seed(lqn.nidx + 1, 0.0); // build-time stage mean
1769 std::vector<std::vector<std::size_t>> rp_member_inbound(lqn.nidx + 1);
1770 std::vector<std::vector<std::size_t>> rp_hop_inbound(lqn.nidx + 1);
1771 std::vector<std::vector<std::size_t>> rp_group_roots;
1772 std::vector<std::size_t> rp_group_reftask;
1773 std::vector<bool> rp_group_head_is_caller;
1774 // classes minted for the path, 0 when absent
1775 std::vector<std::size_t> rp_ref_stage;
1776 std::vector<std::size_t> rp_hop_cls(lqn.nidx + 1, 0), rp_hop_ret(lqn.nidx + 1, 0),
1777 rp_ret_cls(lqn.nidx + 1, 0);
1778 std::vector<std::size_t> rp_gate(lqn.ncalls + 1, 0), rp_aux(lqn.ncalls + 1, 0),
1779 rp_resume(lqn.ncalls + 1, 0);
1780
1781 /** Build-time seed of a hop stage: the entry's own host demand. */
1782 auto host_demand_of_entry = [&](std::size_t eidx_) {
1783 double d = 0.0;
1784 for (std::size_t a_ : lqn.actsof[eidx_]) {
1785 if (lqn.hostdem[a_].disabled) continue;
1786 const double m_ = dbl(lqn.hostdem[a_].mean);
1787 if (std::isfinite(m_)) d += m_;
1788 }
1789 return d;
1790 };
1791 const auto in_idxset = [&](std::size_t t_) {
1792 return std::find(idxSet.begin(), idxSet.end(), t_) != idxSet.end();
1793 };
1794
1795 // Port of resolveRefPath: a pure query over lqn plus the layer's own
1796 // server set, run before any class exists because the first pass needs
1797 // the answer to size the caller populations.
1798 [&]() {
1799 if (interlockMethod != "refpath" || flat || iscachelayer || hasfork || hasjoin) return;
1800 // a Fork flattens every positive alpha in the visit solve; a replicated
1801 // element carries a per-replica njobs convention a shared chain breaks
1802 if (reduce_fanout || nreplicas > 1) return;
1803 for (std::size_t t_ : callers)
1804 if (lqn.repl[t_] > 1.0 || single_replica_tasks.count(t_) > 0) return;
1805 const lqn::LqnRefRoutes<T> R =
1806 lqn::lqn_ref_routes(lqn, callers, opt.interlock_maxpaths, idxSet);
1807 if (!R.why.empty()) {
1808 // named, because the discontinuity is otherwise invisible
1809 std::cerr << "[LINE] Warning: layer '" << lqn.hashnames[idx]
1810 << "' falls back to the 'ilrate' interlock method: " << R.why
1811 << ".\n";
1812 return;
1813 }
1814 std::vector<std::size_t> keep;
1815 for (std::size_t g_ = 0; g_ < R.groups.size(); ++g_) {
1816 const lqn::LqnRefGroup<T>& G = R.groups[g_];
1817 bool wanted = G.members.size() >= 2;
1818 // a lone caller below a reference task: no merge, but the path
1819 // above it becomes explicit instead of a fitted idle time
1820 if (!wanted && opt.interlock_refpath_scope == "all" && !G.head_is_caller &&
1821 !G.members.empty())
1822 wanted = true;
1823 if (!wanted) continue;
1824 // The merge only pays when the chain population it installs is
1825 // STRICTLY below the caller pools it replaces; otherwise the
1826 // group stays on 'ilrate' (17-external-interlock).
1827 double pool = 0.0;
1828 for (std::size_t t_ : G.members) pool += caller_pool(t_);
1829 const double chainPop = mult[G.reftask] * lqn.repl[G.reftask];
1830 if (!(pool > chainPop)) continue;
1831 keep.push_back(g_);
1832 }
1833 if (keep.empty()) return;
1834 // Stage the whole decision, then commit: a refusal found on the
1835 // second group must not leave the first one half-applied.
1836 std::vector<bool> sMember(lqn.nidx + 1, false), sSpliced(lqn.nidx + 1, false),
1837 sHop(lqn.nidx + 1, false), sExpand(lqn.ncalls + 1, false);
1838 std::vector<std::size_t> sChildTask(lqn.ncalls + 1, 0),
1839 sChildEntry(lqn.ncalls + 1, 0);
1840 std::vector<T> sRetShare(lqn.ncalls + 1, Tzero()), sCallMean(lqn.ncalls + 1, Tzero());
1841 std::vector<std::vector<std::size_t>> sHopChildren(lqn.nidx + 1),
1842 sMemberInbound(lqn.nidx + 1), sHopInbound(lqn.nidx + 1);
1843 std::vector<double> sHopSeed(lqn.nidx + 1, 0.0);
1844 std::vector<std::vector<std::size_t>> sRoots;
1845 std::vector<std::size_t> sRefTask;
1846 std::vector<bool> sHeadIsCaller;
1847 for (std::size_t g_ : keep) {
1848 const lqn::LqnRefGroup<T>& G = R.groups[g_];
1849 const std::size_t nE = G.entries.size();
1850 // does this DAG node lead to a caller? reverse topological order
1851 std::vector<bool> leads(G.ismember.begin(), G.ismember.end());
1852 for (std::size_t i_ = nE; i_-- > 0;)
1853 for (const lqn::LqnRefCall<T>& c_ : G.calls)
1854 if (c_.from == i_ && leads[c_.to]) leads[i_] = true;
1855 for (std::size_t t_ : G.members)
1856 if (!caller_needs_class(t_)) return; // no client class to enter
1857 if (G.head_is_caller && !caller_needs_class(G.reftask)) return;
1858 // an infinite reference pool has no chain population
1859 if (!G.head_is_caller && !std::isfinite(lqn.maxmult[G.reftask])) return;
1860 for (std::size_t i_ = 0; i_ < nE; ++i_) {
1861 if (!leads[i_] || G.ismember[i_]) continue;
1862 // a hop served by this layer's own station would be in two places
1863 if (in_idxset(G.etask[i_]) || lqn.repl[G.etask[i_]] > 1.0) return;
1864 }
1865 for (const lqn::LqnRefCall<T>& c_ : G.calls) {
1866 if (!leads[c_.to]) continue;
1867 if (in_idxset(G.etask[c_.to])) return; // callee served by this layer
1868 }
1869 // commit this group into the staging tables
1870 const std::size_t head = G.reftask;
1871 for (std::size_t t_ : G.members) {
1872 sMember[t_] = true;
1873 sSpliced[t_] = !(G.head_is_caller && t_ == head);
1874 }
1875 std::vector<T> inTotMember(lqn.nidx + 1, Tzero()), inTotHop(lqn.nidx + 1, Tzero());
1876 for (const lqn::LqnRefCall<T>& c_ : G.calls) {
1877 if (!leads[c_.to]) continue;
1878 const std::size_t cidx_ = c_.cidx;
1879 const std::size_t fromEidx = G.entries[c_.from];
1880 sExpand[cidx_] = true;
1881 sCallMean[cidx_] = lqn.callproc_mean[cidx_];
1882 sHopChildren[fromEidx].push_back(cidx_);
1883 if (G.ismember[c_.to]) {
1884 const std::size_t tv = G.etask[c_.to];
1885 sChildTask[cidx_] = tv;
1886 sMemberInbound[tv].push_back(cidx_);
1887 inTotMember[tv] = T(inTotMember[tv] + c_.vcall);
1888 } else {
1889 const std::size_t ev = G.entries[c_.to];
1890 sChildEntry[cidx_] = ev;
1891 sHopInbound[ev].push_back(cidx_);
1892 inTotHop[ev] = T(inTotHop[ev] + c_.vcall);
1893 }
1894 sRetShare[cidx_] = c_.vcall; // normalised below
1895 }
1896 for (const lqn::LqnRefCall<T>& c_ : G.calls) {
1897 const std::size_t cidx_ = c_.cidx;
1898 if (!sExpand[cidx_]) continue;
1899 const T tot = sChildTask[cidx_] > 0 ? inTotMember[sChildTask[cidx_]]
1900 : inTotHop[sChildEntry[cidx_]];
1901 sRetShare[cidx_] = dbl(tot) > GlobalConstants::FineTol
1902 ? T(sRetShare[cidx_] / tot)
1903 : Tzero();
1904 }
1905 for (std::size_t i_ = 0; i_ < nE; ++i_)
1906 if (leads[i_] && !G.ismember[i_]) {
1907 const std::size_t ev = G.entries[i_];
1908 sHop[ev] = true;
1909 sHopSeed[ev] = host_demand_of_entry(ev);
1910 }
1911 if (!G.head_is_caller) {
1912 // EVERY entry of the reference task becomes a stage, not only
1913 // those reaching this layer, or the 1/nentries split would
1914 // shrink its denominator and over-drive the chain
1915 for (std::size_t ev : lqn.entriesof[G.reftask])
1916 if (!sHop[ev]) {
1917 sHop[ev] = true;
1918 sHopSeed[ev] = host_demand_of_entry(ev);
1919 }
1920 }
1921 sRoots.push_back(lqn.entriesof[head]);
1922 sRefTask.push_back(head);
1923 sHeadIsCaller.push_back(G.head_is_caller);
1924 }
1925 rp_member = sMember;
1926 rp_spliced = sSpliced;
1927 rp_hop = sHop;
1928 rp_expand = sExpand;
1929 rp_child_task = sChildTask;
1930 rp_child_entry = sChildEntry;
1931 rp_ret_share = sRetShare;
1932 rp_call_mean = sCallMean;
1933 rp_hop_children = sHopChildren;
1934 rp_hop_seed = sHopSeed;
1935 rp_member_inbound = sMemberInbound;
1936 rp_hop_inbound = sHopInbound;
1937 rp_group_roots = sRoots;
1938 rp_group_reftask = sRefTask;
1939 rp_group_head_is_caller = sHeadIsCaller;
1940 rp_on = true;
1941 }();
1942
1943 // ---- first pass: create the classes --------------------------------
1944 for (std::size_t tidx_caller : callers) {
1945 if (caller_needs_class(tidx_caller)) {
1946 double nj = njobs(tidx_caller, idx);
1947 if (nj == 0.0) {
1948 nj = caller_pool(tidx_caller);
1949 njobs(tidx_caller, idx) = nj;
1950 }
1951 // A caller merged into a reference chain brings NO population of
1952 // its own: the customers are the reference task's. Only the CLASS
1953 // is zeroed; njobs still records the pool for update_think_times.
1954 const bool rp_spliced_here = rp_on && rp_spliced[tidx_caller];
1955 qn::JobClass jc;
1956 jc.name = lqn.hashnames[tidx_caller];
1957 jc.type = JobClassType::CLOSED;
1958 jc.population = rp_spliced_here ? 0.0 : nj;
1959 jc.refstat = m.clientIdx;
1960 jc.completes = false;
1961 // ONE reference class per chain: a merged caller yields it to the
1962 // chain head and is normalised through layer_normclass instead
1963 jc.is_ref_class = !rp_spliced_here;
1964 jc.attr_kind = int(LqnElement::TASK);
1965 jc.attr_idx = tidx_caller;
1966 cls[tidx_caller] = m.add_class(jc);
1967 m.attr_tasks.emplace_back(cls[tidx_caller], tidx_caller);
1968 if (lqn.isref[tidx_caller]) {
1969 m.set_service(m.clientIdx, cls[tidx_caller], thinkproc[tidx_caller]);
1970 } else if (rp_spliced_here) {
1971 // the path above this caller is explicit now, so there is no
1972 // idle time left to fit and no update-map row to refresh it
1973 m.set_service(m.clientIdx, cls[tidx_caller], Distrib<T>::immediate());
1974 } else {
1975 // a served task's declared think time is not a per-request
1976 // delay, so the seed carries none either; update_think_times
1977 // replaces this from the first iteration on
1978 m.set_service(m.clientIdx, cls[tidx_caller], Distrib<T>::immediate());
1979 thinkt_map.push_back({idx, tidx_caller, m.clientIdx, cls[tidx_caller]});
1980 }
1981
1982 for (std::size_t eidx : lqn.entriesof[tidx_caller]) {
1983 qn::JobClass ec;
1984 ec.name = lqn.hashnames[eidx];
1985 ec.type = JobClassType::CLOSED;
1986 ec.population = 0.0;
1987 ec.refstat = m.clientIdx;
1988 ec.completes = false;
1989 ec.attr_kind = int(LqnElement::ENTRY);
1990 ec.attr_idx = eidx;
1991 cls[eidx] = m.add_class(ec);
1992 m.attr_entries.emplace_back(cls[eidx], eidx);
1993 m.set_service(m.clientIdx, cls[eidx], Distrib<T>::immediate());
1994 }
1995 }
1996
1997 for (std::size_t aidx : lqn.actsof[tidx_caller]) {
1998 if (caller_acts_visible(tidx_caller)) {
1999 qn::JobClass ac;
2000 ac.name = lqn.hashnames[aidx];
2001 ac.type = JobClassType::CLOSED;
2002 ac.population = 0.0;
2003 ac.refstat = m.clientIdx;
2004 ac.completes = false;
2005 ac.attr_kind = int(LqnElement::ACTIVITY);
2006 ac.attr_idx = aidx;
2007 cls[aidx] = m.add_class(ac);
2008 m.attr_activities.emplace_back(cls[aidx], aidx);
2009 // The host demand is served at the processor's own station
2010 // when this layer holds it; everywhere else the activity is
2011 // a surrogate delay at the client.
2012 const std::size_t hidx = lqn.parent[lqn.parent[aidx]];
2013 if (servers_for(hidx).empty())
2014 m.set_service(m.clientIdx, cls[aidx], servtproc[aidx]);
2015 }
2016 for (std::size_t cidx : lqn.callsof[aidx]) {
2017 if (lqn.calltype[cidx] == CallType::ASYNC) {
2018 // An async call is a stream, not a visit: the caller does
2019 // not block, so it carries no closed class and instead
2020 // drives an open chain of its own out of the Source.
2021 const std::size_t adst = lqn.parent[lqn.callpair_dst[cidx]];
2022 if (servers_for(adst).empty()) continue;
2023 qn::JobClass oc;
2024 oc.name = lqn.callhashnames[cidx];
2025 oc.type = JobClassType::OPEN;
2026 oc.population = std::numeric_limits<double>::infinity();
2027 oc.refstat = sourceStation;
2028 oc.completes = false;
2029 oc.is_ref_class = false;
2030 oc.attr_kind = int(LqnElement::CALL);
2031 oc.attr_idx = cidx;
2032 callcls[cidx] = m.add_class(oc);
2033 m.attr_calls.push_back({callcls[cidx], cidx, lqn.callpair_src[cidx],
2034 lqn.callpair_dst[cidx]});
2035 // A NEGLIGIBLE seed, not the reference's Immediate. The
2036 // rate is replaced from the caller's throughput on the
2037 // first post(), but apply_sink_closure weights the
2038 // Sink->Source arcs by it in between, and Immediate is
2039 // rate 1/0: the closed classes sharing the closure then
2040 // get NaN visits. fj_mmt seeds its own open classes the
2041 // same way, for the same reason.
2042 m.set_service(sourceStation, callcls[cidx],
2044 num_traits<T>::from_double(GlobalConstants::FineTol)));
2045 T minRespTA = Tzero();
2046 for (std::size_t ta : lqn.actsof[adst]) minRespTA += lqn.hostdem[ta].mean;
2047 for (std::size_t st : servers_for(adst)) {
2048 m.set_service(st, callcls[cidx], Distrib<T>::exp_mean(minRespTA));
2049 call_map.push_back({idx, cidx, st, callcls[cidx]});
2050 }
2051 arv_call_map.push_back({idx, cidx, sourceStation, callcls[cidx]});
2052 continue;
2053 }
2054 if (lqn.calltype[cidx] != CallType::SYNC) continue;
2055 const std::size_t gid = group_of_call[cidx];
2056 if (gid != 0) {
2057 // ONE dispatch with n destinations. The members SHARE the
2058 // dispatch class, which is both the class the strategy
2059 // routes and the class that visits the targets: the hop
2060 // must not switch class, because a state-dependent
2061 // routing function is evaluated at zero off the class
2062 // diagonal. The class switch goes on the return arc
2063 // instead, into a group class the job continues in.
2064 if (grp_dispatch[gid] == 0) {
2065 // The strategy is a property of a NODE and routes
2066 // over that node's links, not over one class's arcs,
2067 // so a dedicated Router whose only links are the
2068 // group's targets is the only place where the choice
2069 // is exactly the group's.
2070 const std::string tag =
2071 lqn.hashnames[aidx] + ".Dispatch" + std::to_string(gid);
2072 grp_router[gid] = m.add_node(tag + ".Router", NodeType::Router, true);
2073 qn::JobClass dc;
2074 dc.name = tag;
2075 dc.type = JobClassType::CLOSED;
2076 dc.population = 0.0;
2077 dc.refstat = m.clientIdx;
2078 dc.completes = false;
2079 dc.attr_kind = int(LqnElement::CALL);
2080 dc.attr_idx = cidx;
2081 grp_dispatch[gid] = m.add_class(dc);
2082 m.set_service(m.clientIdx, grp_dispatch[gid], Distrib<T>::immediate());
2083 qn::JobClass gc;
2084 gc.name = lqn.callhashnames[cidx] + ".Group" + std::to_string(gid);
2085 gc.type = JobClassType::CLOSED;
2086 gc.population = 0.0;
2087 gc.refstat = m.clientIdx;
2088 gc.completes = false;
2089 gc.attr_kind = int(LqnElement::CALL);
2090 gc.attr_idx = cidx;
2091 grp_class[gid] = m.add_class(gc);
2092 m.set_service(m.clientIdx, grp_class[gid], Distrib<T>::immediate());
2093 routed_group_sites.push_back(
2094 {grp_router[gid], grp_dispatch[gid], gid});
2095 }
2096 callcls[cidx] = grp_dispatch[gid];
2097 m.attr_calls.push_back({callcls[cidx], cidx, lqn.callpair_src[cidx],
2098 lqn.callpair_dst[cidx]});
2099 for (std::size_t st2 : servers_for(lqn.parent[lqn.callpair_dst[cidx]]))
2100 m.set_service(st2, callcls[cidx], callservtproc[cidx]);
2101 continue;
2102 }
2103 qn::JobClass cc;
2104 cc.name = lqn.callhashnames[cidx];
2105 cc.type = JobClassType::CLOSED;
2106 cc.population = 0.0;
2107 cc.refstat = m.clientIdx;
2108 cc.completes = false;
2109 cc.attr_kind = int(LqnElement::CALL);
2110 cc.attr_idx = cidx;
2111 callcls[cidx] = m.add_class(cc);
2112 m.attr_calls.push_back({callcls[cidx], cidx, lqn.callpair_src[cidx],
2113 lqn.callpair_dst[cidx]});
2114 // An upper bound on the server's response, replaced at the
2115 // first iteration by the measured one. The station seeded is
2116 // the CALLEE's under `flat` and the layer's own under
2117 // `srvn`, where the latter also seeds calls that leave the
2118 // layer -- a phantom rate on a class with no visits here,
2119 // kept because removing it moves the fixed point (see
2120 // seedCallService in buildLayersRecursive.m).
2121 const std::size_t seedidx = flat ? lqn.parent[lqn.callpair_dst[cidx]] : idx;
2122 T minRespT = Tzero();
2123 for (std::size_t ta : lqn.actsof[seedidx]) minRespT += lqn.hostdem[ta].mean;
2124 for (std::size_t st : servers_for(seedidx))
2125 m.set_service(st, callcls[cidx], Distrib<T>::exp_mean(minRespT));
2126 // .Aux class rationale: see _kb/06-solver-catalog.md (cpp port notes: solver_ln.h)
2127 if (lqn.callproc_mean[cidx] !=
2128 num_traits<T>::from_double(double(nreplicas))) {
2129 qn::JobClass xc;
2130 xc.name = lqn.callhashnames[cidx] + ".Aux";
2131 xc.type = JobClassType::CLOSED;
2132 xc.population = 0.0;
2133 xc.refstat = m.clientIdx;
2134 xc.completes = false;
2135 xc.attr_kind = int(LqnElement::CALL);
2136 xc.attr_idx = cidx;
2137 auxcallcls[cidx] = m.add_class(xc);
2138 m.set_service(m.clientIdx, auxcallcls[cidx], Distrib<T>::immediate());
2139 }
2140 }
2141 }
2142 }
2143
2144 // ---- reference-path classes (buildRefPathClasses) ------------------
2145 // Minted with the caller classes. A descent issued from a caller's own
2146 // activity graph REUSES the call classes that already exist as its gate
2147 // and its geometric auxiliary: fresh ones would leave those two with no
2148 // inbound arc at all.
2149 auto rp_class = [&](const std::string& nm, double pop, std::size_t attr) {
2150 qn::JobClass rc;
2151 rc.name = nm;
2152 rc.type = JobClassType::CLOSED;
2153 rc.population = pop;
2154 rc.refstat = m.clientIdx;
2155 rc.completes = false;
2156 rc.is_ref_class = false;
2157 rc.attr_kind = -1;
2158 rc.attr_idx = attr;
2159 const std::size_t k_ = m.add_class(rc);
2160 m.set_service(m.clientIdx, k_, Distrib<T>::immediate());
2161 return k_;
2162 };
2163 if (rp_on) {
2164 for (std::size_t g_ = 0; g_ < rp_group_reftask.size(); ++g_) {
2165 if (rp_group_head_is_caller[g_]) {
2166 rp_ref_stage.push_back(0); // the caller class IS the head
2167 continue;
2168 }
2169 const std::size_t rt = rp_group_reftask[g_];
2170 const std::size_t k_ =
2171 rp_class(lqn.hashnames[rt] + ".RefPath", mult[rt] * lqn.repl[rt], rt);
2172 m.classes[k_ - 1].is_ref_class = true;
2173 m.set_service(m.clientIdx, k_, thinkproc[rt]);
2174 rp_ref_stage.push_back(k_);
2175 }
2176 for (std::size_t ev = 1; ev <= lqn.nidx; ++ev) {
2177 if (!rp_hop[ev]) continue;
2178 const std::size_t k_ = rp_class(lqn.hashnames[ev] + ".RefHop", 0.0, ev);
2179 if (rp_hop_seed[ev] > GlobalConstants::FineTol)
2180 m.set_service(m.clientIdx, k_,
2181 Distrib<T>::exp_mean(num_traits<T>::from_double(rp_hop_seed[ev])));
2182 rp_hop_cls[ev] = k_;
2183 // the entry's own residence, MINUS the descents re-expanded below
2184 // it, PLUS the wait its task's threads queue for
2185 refpath_map.push_back({idx, m.clientIdx, k_, 1, ev, 1.0, 0});
2186 for (std::size_t c_ : rp_hop_children[ev])
2187 refpath_map.push_back({idx, m.clientIdx, k_, 2, c_, -1.0, ev});
2188 refpath_map.push_back({idx, m.clientIdx, k_, 3, lqn.parent[ev], 1.0, 0});
2189 rp_hop_ret[ev] = rp_class(lqn.hashnames[ev] + ".RefHopRet", 0.0, ev);
2190 }
2191 for (std::size_t tv = 1; tv <= lqn.nidx; ++tv)
2192 if (rp_spliced[tv]) rp_ret_cls[tv] = rp_class(lqn.hashnames[tv] + ".RefRet", 0.0, tv);
2193 for (std::size_t cidx_ = 1; cidx_ <= lqn.ncalls; ++cidx_) {
2194 if (!rp_expand[cidx_]) continue;
2195 if (callcls[cidx_] != 0) {
2196 rp_gate[cidx_] = callcls[cidx_];
2197 m.set_service(m.clientIdx, rp_gate[cidx_], Distrib<T>::immediate());
2198 if (auxcallcls[cidx_] != 0) rp_aux[cidx_] = auxcallcls[cidx_];
2199 } else {
2200 rp_gate[cidx_] = rp_class(lqn.callhashnames[cidx_] + ".RefGate", 0.0, cidx_);
2201 if (rp_call_mean[cidx_] != Tone())
2202 rp_aux[cidx_] =
2203 rp_class(lqn.callhashnames[cidx_] + ".RefGate.Aux", 0.0, cidx_);
2204 }
2205 // A MEMBER entered through this gate has its subgraph explicit
2206 // here; the one cost it does not show is the wait for one of its
2207 // own threads, which lives in its own task layer.
2208 if (rp_child_task[cidx_] > 0)
2209 refpath_map.push_back(
2210 {idx, m.clientIdx, rp_gate[cidx_], 3, rp_child_task[cidx_], 1.0, 0});
2211 rp_resume[cidx_] = rp_class(lqn.callhashnames[cidx_] + ".RefResume", 0.0, cidx_);
2212 }
2213 }
2214
2215 // ---- second pass: routing out of the entries -------------------------
2216 struct Ctx {
2217 std::size_t curclass;
2218 int jobpos; // 1 = at the client, 2 = at a server replica, 3 = at the Cache
2219 // The nodes the job stands at when jobpos is atServer. Under `srvn`
2220 // these are always the layer's own server replicas; under `flat`
2221 // they are whichever element's stations the last hop landed on, so
2222 // the position has to be carried rather than assumed.
2223 std::vector<std::size_t> curnodes;
2224 };
2225 const int atClient = 1, atServer = 2, atCache = 3;
2226 std::vector<int> jobposkey(lqn.nidx + 1, atClient);
2227 std::vector<std::size_t> curclasskey(lqn.nidx + 1, 0);
2228 std::vector<std::vector<std::size_t>> curnodeskey(lqn.nidx + 1);
2229
2230 // Port of routeRefPathDescent: one explicit descent into the callee of
2231 // CIDX_, replacing the surrogate delay that would charge its whole
2232 // residence at the client. The multiplicity is a geometric loop per
2233 // edge, never one flat loop of the product, and the m < 1 branch is not
2234 // optional: 1 - 1/m goes negative below one.
2235 auto rp_descent = [&](std::size_t cidx_, Ctx st) -> Ctx {
2236 const std::size_t gate = rp_gate[cidx_], res = rp_resume[cidx_], aux = rp_aux[cidx_];
2237 const T m_ = rp_call_mean[cidx_];
2238 const std::size_t enter = rp_child_task[cidx_] > 0 ? cls[rp_child_task[cidx_]]
2239 : rp_hop_cls[rp_child_entry[cidx_]];
2240 if (st.jobpos == atClient)
2241 m.set_route(st.curclass, gate, clientNode, clientNode, Tone());
2242 else
2243 for (std::size_t nd : st.curnodes) m.set_route(st.curclass, gate, nd, clientNode, Tone());
2244 std::size_t cont;
2245 if (m_ == Tone()) {
2246 m.set_route(gate, enter, clientNode, clientNode, Tone());
2247 cont = res;
2248 } else if (m_ > Tone()) {
2249 m.set_route(gate, enter, clientNode, clientNode, Tone());
2250 m.set_route(res, enter, clientNode, clientNode, T(Tone() - Tone() / m_));
2251 m.set_route(res, aux, clientNode, clientNode, T(Tone() / m_));
2252 cont = aux;
2253 } else {
2254 m.set_route(gate, enter, clientNode, clientNode, m_);
2255 m.set_route(gate, aux, clientNode, clientNode, T(Tone() - m_));
2256 m.set_route(res, aux, clientNode, clientNode, Tone());
2257 cont = aux;
2258 }
2259 st.jobpos = atClient;
2260 st.curnodes.clear();
2261 st.curclass = cont;
2262 return st;
2263 };
2264
2265 std::function<Ctx(std::size_t, std::size_t, Ctx)> recur =
2266 [&](std::size_t tidx_caller, std::size_t aidx, Ctx st) -> Ctx {
2267 jobposkey[aidx] = st.jobpos;
2268 curclasskey[aidx] = st.curclass;
2269 curnodeskey[aidx] = st.curnodes;
2270 const std::vector<std::size_t> nexts = lqn.graph.succ(aidx);
2271 std::size_t lastEntryClass = st.curclass;
2272 // fork pre-state save/restore rationale: see _kb/06-solver-catalog.md (cpp port notes: solver_ln.h)
2273 bool next_is_fork = false;
2274 for (std::size_t sx : nexts)
2275 if (lqn.actposttype[sx] == PrecedenceType::POST_AND) next_is_fork = true;
2276 const Ctx preFork = st;
2277 std::vector<std::size_t> andSuccs;
2278 for (std::size_t sx : nexts)
2279 if (lqn.actposttype[sx] == PrecedenceType::POST_AND) andSuccs.push_back(sx);
2280
2281 for (std::size_t k = 0; k < nexts.size(); ++k) {
2282 const std::size_t nextaidx = nexts[k];
2283 if (next_is_fork) st = preFork;
2284 bool isLoop = lqn.graph.get(aidx, nextaidx) != lqn.dag.get(aidx, nextaidx);
2285 if (lqn.parent[aidx] != lqn.parent[nextaidx]) {
2286 // a call to an entry of another task
2287 std::size_t cidx = 0;
2288 for (std::size_t c : lqn.callsof[aidx])
2289 if (lqn.callpair_dst[c] == nextaidx) cidx = c;
2290 if (cidx == 0) continue;
2291 // FWD lays down no routing: its work is already carried by the pseudo SYNC arc
2292 if (lqn.calltype[cidx] != CallType::SYNC) continue;
2293 const std::size_t gid = group_of_call[cidx];
2294 if (gid != 0) {
2295 // The whole group is routed once, at its FIRST member;
2296 // the others are the same dispatch and lay down nothing.
2297 if (group_members[gid].empty() || group_members[gid][0] != cidx) continue;
2298 // The job is switched into the dispatch class while still
2299 // at the client, so (client, dispatch) carries exactly
2300 // the n arcs the strategy chooses among. The 1/n split
2301 // laid down here is the probabilistic reading a solver
2302 // without state-dependent routing would see; the
2303 // declared strategy replaces it below.
2304 std::vector<std::size_t> tnode2, tstat2, tcall2;
2305 for (std::size_t mc : group_members[gid]) {
2306 const std::size_t tt = lqn.parent[lqn.callpair_dst[mc]];
2307 if (servers_for(tt).empty()) continue;
2308 tnode2.push_back(server_nodes_for(tt)[0]);
2309 tstat2.push_back(servers_for(tt)[0]);
2310 tcall2.push_back(mc);
2311 }
2312 if (tnode2.size() < 2) continue; // not enough of it is here
2313 const std::size_t fromNode2 =
2314 st.jobpos == atClient ? clientNode : st.curnodes[0];
2315 const std::size_t dc = grp_dispatch[gid], gc = grp_class[gid];
2316 m.set_route(st.curclass, dc, fromNode2, grp_router[gid], Tone());
2317 const T share2 =
2318 T(Tone() / num_traits<T>::from_int(int(tnode2.size())));
2319 for (std::size_t d = 0; d < tnode2.size(); ++d) {
2320 m.set_route(dc, dc, grp_router[gid], tnode2[d], share2);
2321 m.set_route(dc, gc, tnode2[d], clientNode, Tone());
2322 m.set_service(tstat2[d], dc, callservtproc[tcall2[d]]);
2323 call_map.push_back({idx, tcall2[d], tstat2[d], dc});
2324 }
2325 st.curclass = gc;
2326 st.jobpos = atClient;
2327 st.curnodes.clear();
2328 continue;
2329 }
2330 if (rp_on && rp_expand[cidx]) {
2331 // the callee is on the reference path into this layer: the
2332 // customer descends into it rather than waiting a delay out
2333 st = rp_descent(cidx, st);
2334 continue;
2335 }
2336 const std::size_t ctgt = lqn.parent[lqn.callpair_dst[cidx]];
2337 st = route_sync_call(m, idx, cidx, st, server_nodes_for(ctgt),
2338 servers_for(ctgt), callcls, auxcallcls, clientNode,
2339 atClient, atServer, flat);
2340 continue;
2341 }
2342 // a successor inside the same task
2343 bool any_entry_succ = false;
2344 for (std::size_t sx : nexts)
2345 if (sx > lqn.eshift && sx <= lqn.eshift + lqn.nentries) any_entry_succ = true;
2346 if (!any_entry_succ) {
2347 st.jobpos = jobposkey[aidx];
2348 st.curclass = curclasskey[aidx];
2349 st.curnodes = curnodeskey[aidx];
2350 } else {
2351 if (k > 0 && nexts[k - 1] > lqn.eshift && nexts[k - 1] <= lqn.eshift + lqn.nentries)
2352 lastEntryClass = st.curclass;
2353 st.jobpos = atClient;
2354 st.curclass = lastEntryClass;
2355 st.curnodes.clear();
2356 }
2357 const T w = lqn.graph.get(aidx, nextaidx);
2358
2359 // THE CACHE READ, buildLayersRecursive.m:743-765. `aidx` is the
2360 // item entry and `nextaidx` the activity bound to it: the job
2361 // goes from the client to the Cache node, which decides the hit
2362 // or the miss and switches the class accordingly, so the read
2363 // class itself is never served anywhere and the branch classes
2364 // pick the work up at the server.
2365 if (iscachelayer && lqn.nitems[aidx] > 0 && cls[nextaidx] != 0) {
2366 const std::size_t readcls = cls[nextaidx];
2367 m.set_route(st.curclass, readcls, clientNode, cacheNode, w);
2368 if (cachepar.pread.size() < m.classes.size())
2369 cachepar.pread.resize(m.classes.size());
2370 cachepar.pread[readcls - 1] = lqn.itemproc[aidx];
2371 const std::vector<std::size_t> hm = lqn.graph.succ(nextaidx);
2372 if (hm.size() != 2)
2373 throw InputError("SolverLN: the cache read '" + lqn.names[nextaidx] +
2374 "' needs exactly one hit and one miss successor");
2375 if (cachepar.hitclass.size() < m.classes.size()) {
2376 cachepar.hitclass.resize(m.classes.size(), 0);
2377 cachepar.missclass.resize(m.classes.size(), 0);
2378 }
2379 cachepar.hitclass[readcls - 1] = cls[hm[0]];
2380 cachepar.missclass[readcls - 1] = cls[hm[1]];
2381 st.jobpos = atCache;
2382 st.curclass = readcls;
2383 st.curnodes.clear();
2384 st = recur(tidx_caller, nextaidx, st);
2385 continue;
2386 }
2387
2388 const bool is_and_join_tail = lqn.actpretype[aidx] == PrecedenceType::PRE_AND;
2389 // the branch index of this successor among the fork's outputs
2390 std::size_t fbranch = 0;
2391 if (next_is_fork)
2392 for (std::size_t q = 0; q < andSuccs.size(); ++q)
2393 if (andSuccs[q] == nextaidx) fbranch = q + 1;
2394
2395 // Where the successor's HOST DEMAND is served: at the station of
2396 // its own processor when this layer holds it, at the client
2397 // otherwise. Under `srvn` that is the host layer's own server
2398 // and nothing else, so this reduces to the `ishostlayer` test
2399 // it replaces.
2400 const std::size_t hidxOf = lqn.parent[lqn.parent[nextaidx]];
2401 const std::vector<std::size_t>& actStations = servers_for(hidxOf);
2402 const std::vector<std::size_t>& actNodes = server_nodes_for(hidxOf);
2403 const bool actAtServer = !actStations.empty();
2404 const std::size_t from = st.jobpos == atClient ? clientNode
2405 : st.jobpos == atCache ? cacheNode
2406 : st.curnodes[0];
2407 // the node a job continues to after this successor is served
2408 for (std::size_t r = 0; r < nreplicas; ++r) {
2409 const std::size_t fromNode =
2410 st.jobpos == atClient ? clientNode
2411 : st.jobpos == atCache ? cacheNode
2412 : st.curnodes[std::min(r, st.curnodes.size() - 1)];
2413 const std::size_t toNode = actAtServer ? actNodes[r] : clientNode;
2414 if (next_is_fork && fbranch > 0) {
2415 m.set_route(st.curclass, st.curclass, fromNode, forkNode, Tone());
2416 if (r == 0) forkClassStack.push_back(st.curclass);
2417 m.set_route(st.curclass, st.curclass, forkNode, forkRouter[fbranch - 1],
2418 Tone());
2419 m.set_route(st.curclass, cls[nextaidx], forkRouter[fbranch - 1], toNode,
2420 Tone());
2421 } else if (is_and_join_tail) {
2422 // rejoin the class the branch was forked from, then
2423 // leave the Join in the successor's class
2424 if (forkClassStack.empty())
2425 throw InputError("SolverLN: an AND join has no matching fork in '" +
2426 lqn.names[aidx] + "'");
2427 const std::size_t forkClass = forkClassStack.back();
2428 if (r + 1 == nreplicas) forkClassStack.pop_back();
2429 m.set_route(st.curclass, forkClass, fromNode, joinNode, Tone());
2430 m.set_route(forkClass, cls[nextaidx], joinNode, toNode, Tone());
2431 } else {
2432 m.set_route(st.curclass, cls[nextaidx], fromNode, toNode, w);
2433 }
2434 if (actAtServer)
2435 m.set_service(actStations[r], cls[nextaidx], lqn.hostdem[nextaidx]);
2436 }
2437 (void)from;
2438 if (actAtServer) {
2439 st.jobpos = atServer;
2440 st.curclass = cls[nextaidx];
2441 st.curnodes = actNodes;
2442 servt_map.push_back({idx, nextaidx, actStations[0], cls[nextaidx]});
2443 } else {
2444 st.jobpos = atClient;
2445 st.curclass = cls[nextaidx];
2446 st.curnodes.clear();
2447 m.set_service(m.clientIdx, cls[nextaidx], servtproc[nextaidx]);
2448 thinkt_map.push_back({idx, nextaidx, m.clientIdx, cls[nextaidx]});
2449 }
2450 if (aidx != nextaidx && !isLoop) {
2451 st = recur(tidx_caller, nextaidx, st);
2452 // close the branch with a reply back to the caller's class; a
2453 // merged caller replies to its OWN return class instead, which
2454 // hands the customer back to whichever descent invoked it
2455 const std::size_t reply =
2456 (rp_on && rp_spliced[tidx_caller]) ? rp_ret_cls[tidx_caller]
2457 : cls[tidx_caller];
2458 if (st.jobpos == atClient) {
2459 m.set_route(st.curclass, reply, clientNode, clientNode, Tone());
2460 } else {
2461 for (std::size_t nd : st.curnodes)
2462 m.set_route(st.curclass, reply, nd, clientNode, Tone());
2463 }
2464 // .Aux completion-guard rationale: see _kb/06-solver-catalog.md (cpp port notes: solver_ln.h)
2465 if (!is_aux_class(m.classes[st.curclass - 1].name))
2466 m.classes[st.curclass - 1].completes = true;
2467 }
2468 }
2469 return st;
2470 };
2471
2472 for (std::size_t tidx_caller : callers) {
2473 if (!caller_needs_class(tidx_caller)) continue;
2474 const std::vector<std::size_t>& ents = lqn.entriesof[tidx_caller];
2475 const T share = T(Tone() / num_traits<T>::from_int(int(ents.size())));
2476 for (std::size_t eidx : ents) {
2477 m.set_route(cls[tidx_caller], cls[eidx], clientNode, clientNode, share);
2478 if (ents.size() > 1)
2479 route_map.push_back({idx, tidx_caller, eidx, m.clientIdx, m.clientIdx,
2480 cls[tidx_caller], cls[eidx]});
2481 Ctx st{cls[eidx], atClient, {}};
2482 recur(tidx_caller, eidx, st);
2483 }
2484 }
2485
2486 // ---- reference-path arcs no activity graph of this layer writes ------
2487 // routeRefPath: the reference stage, the hop stages and every callee's
2488 // return split.
2489 if (rp_on) {
2490 for (std::size_t g_ = 0; g_ < rp_group_reftask.size(); ++g_) {
2491 if (rp_group_head_is_caller[g_]) continue; // its own graph carries the descent
2492 std::vector<std::size_t> roots;
2493 for (std::size_t r_ : rp_group_roots[g_])
2494 if (rp_hop[r_]) roots.push_back(r_);
2495 if (roots.empty()) continue;
2496 // the REF task picks among its entries: a PROBABILISTIC split
2497 const T share = T(Tone() / num_traits<T>::from_int(int(roots.size())));
2498 for (std::size_t r_ : roots) {
2499 m.set_route(rp_ref_stage[g_], rp_hop_cls[r_], clientNode, clientNode, share);
2500 m.set_route(rp_hop_ret[r_], rp_ref_stage[g_], clientNode, clientNode, Tone());
2501 }
2502 }
2503 // the descents below a hop are a SERIES: one cycle traverses all of them
2504 for (std::size_t ev = 1; ev <= lqn.nidx; ++ev) {
2505 if (!rp_hop[ev]) continue;
2506 Ctx cur{rp_hop_cls[ev], atClient, {}};
2507 for (std::size_t c_ : rp_hop_children[ev]) cur = rp_descent(c_, cur);
2508 m.set_route(cur.curclass, rp_hop_ret[ev], clientNode, clientNode, Tone());
2509 }
2510 // Return arcs, split by the share of the callee's invocations each
2511 // descent contributes: v(p_i)*m_i, exact rather than a heuristic.
2512 for (std::size_t tv = 1; tv <= lqn.nidx; ++tv) {
2513 if (!rp_spliced[tv]) continue;
2514 for (std::size_t c_ : rp_member_inbound[tv])
2515 if (rp_ret_share[c_] > Tzero())
2516 m.set_route(rp_ret_cls[tv], rp_resume[c_], clientNode, clientNode,
2517 rp_ret_share[c_]);
2518 }
2519 for (std::size_t ev = 1; ev <= lqn.nidx; ++ev) {
2520 if (!rp_hop[ev]) continue;
2521 for (std::size_t c_ : rp_hop_inbound[ev])
2522 if (rp_ret_share[c_] > Tzero())
2523 m.set_route(rp_hop_ret[ev], rp_resume[c_], clientNode, clientNode,
2524 rp_ret_share[c_]);
2525 }
2526 }
2527
2528 // ---- open streams, laid down AFTER the activity-graph walk ----------
2529 // buildLayersRecursive.m:473 puts the entry-arrival routing here for the
2530 // same reason: the walk above rewrites whole rows of P and would erase
2531 // an arc written before it. Both kinds route Source -> server -> Sink
2532 // within one class, so they never join a closed class in a chain --
2533 // refresh_chains would reject that, the two carrying different refstats.
2534 if (sourceStation != 0) {
2535 const T nrep = num_traits<T>::from_int(int(nreplicas));
2536 for (std::size_t cidx : async_here) {
2537 const std::size_t oc = callcls[cidx];
2538 if (oc == 0) continue;
2539 // the stream enters the CALLEE's station, which under `srvn` is
2540 // this layer's own server and under `flat` one among many
2541 const std::vector<std::size_t>& anode =
2542 server_nodes_for(lqn.parent[lqn.callpair_dst[cidx]]);
2543 const T callmean = lqn.callproc_mean[cidx];
2544 if (callmean < Tone()) {
2545 // fewer than one call per firing: a single Bernoulli pass
2546 m.set_route(oc, oc, sourceNode, sinkNode, T(Tone() - callmean));
2547 for (std::size_t r = 0; r < anode.size(); ++r) {
2548 m.set_route(oc, oc, sourceNode, anode[r], T(callmean / nrep));
2549 m.set_route(oc, oc, anode[r], sinkNode, Tone());
2550 }
2551 } else {
2552 // callmean visits in expectation, as a geometric self-loop
2553 const T p = T(Tone() / callmean);
2554 for (std::size_t r = 0; r < anode.size(); ++r) {
2555 m.set_route(oc, oc, sourceNode, anode[r], T(Tone() / nrep));
2556 for (std::size_t q = 0; q < anode.size(); ++q)
2557 m.set_route(oc, oc, anode[r], anode[q], T((Tone() - p) / nrep));
2558 m.set_route(oc, oc, anode[r], sinkNode, p);
2559 }
2560 }
2561 }
2562 for (std::size_t eidx : open_entries) {
2563 qn::JobClass eo;
2564 eo.name = lqn.hashnames[eidx] + "_Open";
2565 eo.type = JobClassType::OPEN;
2566 eo.population = std::numeric_limits<double>::infinity();
2567 eo.refstat = sourceStation;
2568 eo.completes = false;
2569 eo.is_ref_class = false;
2570 eo.attr_kind = int(LqnElement::ENTRY);
2571 eo.attr_idx = eidx;
2572 const std::size_t ec = m.add_class(eo);
2573 // the arrival is exogenous and fixed, so it is never reseeded
2574 m.set_service(sourceStation, ec, lqn.arrival[eidx]);
2575 // entries are Immediate; the work is the activity bound to them
2576 std::size_t bound = 0;
2577 for (std::size_t sx : lqn.graph.succ(eidx))
2578 if (bound == 0 && sx > lqn.ashift) bound = sx;
2579 const Distrib<T>& svc = bound != 0 ? servtproc[bound] : servtproc[eidx];
2580 // the arrival enters the processor of the entry's task under
2581 // host layering, and the task's own station under `flat`
2582 const std::vector<std::size_t>& ostat =
2583 flat ? servers_for(lqn.parent[eidx]) : server;
2584 const std::vector<std::size_t>& onode =
2585 flat ? server_nodes_for(lqn.parent[eidx]) : serverNode;
2586 for (std::size_t r = 0; r < ostat.size(); ++r) {
2587 m.set_service(ostat[r], ec, svc);
2588 m.set_route(ec, ec, sourceNode, onode[r], T(Tone() / nrep));
2589 m.set_route(ec, ec, onode[r], sinkNode, Tone());
2590 }
2591 }
2592 }
2593
2594 // ---- admission constraint on the server station ---------------------
2595 // buildLayersRecursive.m:596-625. The constraint is stated over LQN
2596 // elements and has to be re-expressed in the layer's own classes: a
2597 // task occupies its host through the classes of its ACTIVITIES, an
2598 // entry is occupied by the classes of the CALLS that target it. The
2599 // columns are summed into the class, never assigned, because several
2600 // classes can stand for one column.
2601 // Rows of the admission constraint, in the layer's own classes: the
2602 // declared lincon first, then the refpath pool rows.
2603 std::vector<std::vector<T>> Arows;
2604 std::vector<T> brows;
2605 if (idx < lqn.lincon_A.size() && lqn.lincon_A[idx].rows() > 0) {
2606 const Matrix<T>& Aelem = lqn.lincon_A[idx];
2607 Matrix<T> Alayer(Aelem.rows(), m.classes.size(), Tzero());
2608 const std::vector<std::size_t>& constrained =
2609 ishostlayer ? lqn.tasksof[idx] : lqn.entriesof[idx];
2610 for (std::size_t j = 0; j < constrained.size() && j < Aelem.cols(); ++j) {
2611 std::vector<std::size_t> layerClasses;
2612 if (ishostlayer) {
2613 for (std::size_t a : lqn.actsof[constrained[j]])
2614 if (cls[a] != 0) layerClasses.push_back(cls[a]);
2615 } else {
2616 for (std::size_t c = 1; c <= lqn.ncalls; ++c)
2617 if (lqn.callpair_dst[c] == constrained[j] && callcls[c] != 0)
2618 layerClasses.push_back(callcls[c]);
2619 }
2620 for (std::size_t k = 0; k < layerClasses.size(); ++k)
2621 for (std::size_t rr = 0; rr < Aelem.rows(); ++rr)
2622 Alayer(rr, layerClasses[k] - 1) =
2623 T(Alayer(rr, layerClasses[k] - 1) + Aelem(rr, j));
2624 }
2625 for (std::size_t rr = 0; rr < Alayer.rows(); ++rr) {
2626 std::vector<T> row(m.classes.size(), Tzero());
2627 for (std::size_t k = 0; k < m.classes.size(); ++k) row[k] = Alayer(rr, k);
2628 Arows.push_back(row);
2629 brows.push_back(lqn.lincon_b[idx][rr]);
2630 }
2631 }
2632 // A merged caller no longer carries its thread pool as a population, so
2633 // the pool is declared as an admission bound on the server instead, BUT
2634 // ONLY WHEN THE LAYER SOLVER CAN HONOUR ONE. Under the default MVA layer
2635 // solver the pool's other carrier, the thread-acquisition wait charged
2636 // at every descent gate and hop stage, carries it alone; emitting the
2637 // region there anyway would change which engine solves the layer.
2638 if (rp_on && layer_solver_declares_region()) {
2639 std::vector<std::size_t> srvEntries;
2640 for (std::size_t sidx : idxSet)
2641 for (std::size_t e : lqn.entriesof[sidx]) srvEntries.push_back(e);
2642 for (std::size_t tv = 1; tv <= lqn.nidx; ++tv) {
2643 if (!rp_spliced[tv]) continue;
2644 const double cap = lqn.maxmult[tv];
2645 if (!std::isfinite(cap) || cap <= 0.0) continue;
2646 std::vector<T> row(m.classes.size(), Tzero());
2647 bool anyc = false;
2648 if (ishostlayer) {
2649 // a task occupies the host through the classes of its activities
2650 for (std::size_t a_ : lqn.actsof[tv])
2651 if (cls[a_] != 0) {
2652 row[cls[a_] - 1] = Tone();
2653 anyc = true;
2654 }
2655 } else {
2656 // a task occupies the server through the calls it makes into it
2657 for (std::size_t a_ : lqn.actsof[tv])
2658 for (std::size_t c_ : lqn.callsof[a_])
2659 if (std::find(srvEntries.begin(), srvEntries.end(),
2660 lqn.callpair_dst[c_]) != srvEntries.end() &&
2661 callcls[c_] != 0) {
2662 row[callcls[c_] - 1] = Tone();
2663 anyc = true;
2664 }
2665 }
2666 if (anyc) {
2667 Arows.push_back(row);
2668 brows.push_back(num_traits<T>::from_double(cap));
2669 }
2670 }
2671 }
2672 {
2673 bool any = false;
2674 for (const std::vector<T>& row : Arows)
2675 for (const T& v : row)
2676 if (v != Tzero()) any = true;
2677 if (any) {
2678 Matrix<T> Alayer(Arows.size(), m.classes.size(), Tzero());
2679 for (std::size_t rr = 0; rr < Arows.size(); ++rr)
2680 for (std::size_t k = 0; k < m.classes.size(); ++k) Alayer(rr, k) = Arows[rr][k];
2681 // ONE region over every replica, not one each: the constraint
2682 // models a passive resource of the server as a whole (a
2683 // semaphore, a connection pool), so the replicas share tokens.
2684 typename qn::NetworkStruct<T>::Region rg;
2685 const std::size_t M = m.stations.size(), K = m.classes.size();
2686 rg.cap.assign(M, std::vector<double>(K + 1, -1.0));
2687 rg.maxmem.assign(M, -1.0);
2688 rg.members.assign(M, false);
2689 rg.rule.assign(K, lang::DropStrategy::WAITQ);
2690 rg.weight.assign(K, Tone());
2691 rg.size.assign(K, Tone());
2692 for (std::size_t r = 0; r < nreplicas; ++r) rg.members[server[r] - 1] = true;
2693 rg.lincon_A = Alayer;
2694 rg.lincon_b = brows;
2695 m.regions.push_back(rg);
2696 }
2697 }
2698
2699 // The Cache node's parameters are only complete once the walk has seen
2700 // every read: pread, hitclass and missclass are all per READ CLASS, and
2701 // the classes are created as the activity graph is traversed.
2702 if (cacheNode != 0) {
2703 cachepar.pread.resize(m.classes.size());
2704 cachepar.hitclass.resize(m.classes.size(), 0);
2705 cachepar.missclass.resize(m.classes.size(), 0);
2706 m.nodeparam[cacheNode] = cachepar;
2707 }
2708
2709 // The declared dispatch replaces the probabilistic split on the router,
2710 // and ONLY on the (router, dispatch class) pair whose arcs are exactly
2711 // the group's targets. The split stays in P underneath, which is what
2712 // `refresh_routing` re-expands for anything that reads a matrix; the SSA
2713 // engine walks the arcs itself and honours the strategy.
2714 for (const std::array<std::size_t, 3>& site : routed_group_sites) {
2715 qn::NodeDef& nd = m.nodes[site[0] - 1];
2716 if (nd.routing.size() < m.classes.size())
2717 nd.routing.resize(m.classes.size(), RoutingStrategy::PROB);
2718 nd.routing[site[1] - 1] = lqn.callgroups[site[2] - 1].strategy;
2719 }
2720
2721 // ---- queue-dependent service rates on the server station ------------
2722 // buildLayersRecursive.m:700-762. Declared over the server's OPERANDS
2723 // (the tasks of a host, the entries of a task) and re-expressed in the
2724 // layer's own classes, on the same mapping the admission constraint
2725 // uses above. The handles are evaluated in two index spaces -- CTMC and
2726 // the exact recursions pass a per-class vector, the AMVA and NC chain
2727 // recursions a per-chain one -- so each carries both column lists and
2728 // picks by the length of what it is handed. The chain list is only known
2729 // after refresh_chains, hence the shared slot filled just below.
2730 std::vector<std::pair<std::shared_ptr<std::vector<std::vector<std::size_t>>>,
2731 std::vector<std::vector<std::size_t>>>> deferred_chaincols;
2732 for (std::size_t sidx : idxSet) {
2733 const bool hasld = sidx < lqn.lldscaling.size() && !lqn.lldscaling[sidx].empty();
2734 const bool hascd = sidx < lqn.cdscaling.size() && bool(lqn.cdscaling[sidx]);
2735 const bool hasjd = sidx < lqn.jdscaling.size() && bool(lqn.jdscaling[sidx]);
2736 const bool haspools = sidx < lqn.pools.size() && !lqn.pools[sidx].empty();
2737 if (!(hasld || hascd || hasjd || haspools)) continue;
2738 const bool sishost = sidx <= lqn.nhosts;
2739 const std::vector<std::size_t>& operandIdx =
2740 sishost ? lqn.tasksof[sidx] : lqn.entriesof[sidx];
2741 std::vector<std::vector<std::size_t>> cols(operandIdx.size());
2742 for (std::size_t j = 0; j < operandIdx.size(); ++j) {
2743 if (sishost) {
2744 for (std::size_t a : lqn.actsof[operandIdx[j]])
2745 if (cls[a] != 0) cols[j].push_back(cls[a] - 1);
2746 } else {
2747 for (std::size_t c = 1; c <= lqn.ncalls; ++c)
2748 if (lqn.callpair_dst[c] == operandIdx[j] && callcls[c] != 0)
2749 cols[j].push_back(callcls[c] - 1);
2750 }
2751 }
2752 const std::size_t R = m.classes.size();
2753 std::shared_ptr<std::vector<std::vector<std::size_t>>> chaincols =
2754 std::make_shared<std::vector<std::vector<std::size_t>>>();
2755 deferred_chaincols.push_back(std::make_pair(chaincols, cols));
2756 for (std::size_t r = 0; r < nreplicas; ++r) {
2757 // The scalings are written straight onto the Station, as the
2758 // region block above writes m.regions: a Layer is a
2759 // NetworkStruct, not the NetworkBuilder that carries the
2760 // set_*_dependence helpers.
2761 qn::Station<T>& stn = m.stations[srv[sidx][r] - 1];
2762 if (hasld) stn.lldscaling = lqn.lldscaling[sidx];
2763 if (hascd) {
2764 // beta_{i,r} is product form only while an operand maps to a
2765 // single class; where it aggregates several, the same
2766 // scaling is emitted as a joint dependence, numerically
2767 // identical but no longer exact.
2768 bool one_class_each = true;
2769 for (std::size_t j = 0; j < cols.size(); ++j)
2770 if (cols[j].size() > 1) one_class_each = false;
2771 const lang::CdScaling<T> h =
2772 layer_dep_handle(lqn.cdscaling[sidx], cols, chaincols, R);
2773 const std::vector<T> pk = layer_peak(lqn.cdscalingpeak[sidx], cols, R);
2774 if (one_class_each) {
2775 stn.cdscaling = h;
2776 stn.cdscalingpeak = pk;
2777 } else {
2778 stn.jdscaling = h;
2779 stn.jdscalingpeak = pk;
2780 }
2781 }
2782 if (hasjd) {
2783 stn.jdscaling = layer_dep_handle(lqn.jdscaling[sidx], cols, chaincols, R);
2784 stn.jdscalingpeak = layer_peak(lqn.jdscalingpeak[sidx], cols, R);
2785 }
2786 if (haspools) {
2787 // A compatibility declaration IS a rate law: the pools clear
2788 // the activated-server rate of api::sn_compat_rate, order
2789 // independent at every integer state. sn_compat_scaling
2790 // normalises it against the rate the SAME population would
2791 // get under full compatibility, so eta isolates the
2792 // compatibility GRAPH and a fully-compatible pool is the
2793 // neutral eta == 1; the low-occupancy loss stays with the
2794 // solver's own multiserver term.
2795 const lqn::ServerPools<T>& pl = lqn.pools[sidx];
2796 const Matrix<T> compat = pl.compat;
2797 const std::vector<double> counts = pl.counts;
2798 const std::vector<T> rates = pl.rates;
2799 lang::CdScaling<T> etaPool = [compat, counts,
2800 rates](const std::vector<T>& nop) {
2801 return std::vector<T>(
2802 1, api::sn_compat_scaling(compat, counts, rates, nop));
2803 };
2804 stn.jdscaling = layer_dep_handle(etaPool, cols, chaincols, R);
2805 stn.jdscalingpeak = layer_peak(
2806 std::vector<T>(cols.size(), num_traits<T>::from_int(1)), cols, R);
2807 }
2808 }
2809 }
2810
2811 if (rp_on) {
2812 // Per-class normalisation task, read by layer_refclass. A merged chain
2813 // holds more than one caller and only ONE reference class, so without
2814 // this every merged caller's residence would come out scaled by the
2815 // product of the call multiplicities along the path.
2816 std::vector<std::size_t> nc(m.classes.size(), 0);
2817 for (std::size_t tv = 1; tv <= lqn.nidx; ++tv) {
2818 if (!rp_member[tv] || cls[tv] == 0) continue;
2819 const std::size_t own = cls[tv];
2820 std::vector<std::size_t> owned{cls[tv]};
2821 for (std::size_t e_ : lqn.entriesof[tv]) owned.push_back(cls[e_]);
2822 for (std::size_t a_ : lqn.actsof[tv]) {
2823 owned.push_back(cls[a_]);
2824 for (std::size_t c_ : lqn.callsof[a_]) {
2825 owned.push_back(callcls[c_]);
2826 owned.push_back(auxcallcls[c_]);
2827 }
2828 }
2829 // an open class (an async call) has no chain reference rate of its own
2830 for (std::size_t q_ : owned)
2831 if (q_ != 0 && m.classes[q_ - 1].type == JobClassType::CLOSED) nc[q_ - 1] = own;
2832 }
2833 layer_normclass[idx] = nc;
2834 ref_path_audit(m, idx, callers, caller_needs_class, rp_spliced, rp_member, rp_hop,
2835 rp_expand, rp_group_head_is_caller, rp_ref_stage, rp_hop_cls,
2836 rp_hop_ret, rp_ret_cls, rp_gate, rp_aux, rp_resume, cls);
2837 }
2838
2839 // a priority-scheduling processor orders its callers by task priority
2840 apply_host_task_priorities(m, idxSet);
2841 m.refresh_chains();
2842 // The chain columns of every handle emitted above, now that the chains
2843 // exist. The slots are shared with the handles, so filling them here
2844 // reaches every replica without rebuilding a single lambda.
2845 for (std::size_t d = 0; d < deferred_chaincols.size(); ++d) {
2846 const std::vector<std::vector<std::size_t>>& cols = deferred_chaincols[d].second;
2847 std::vector<std::vector<std::size_t>>& out = *deferred_chaincols[d].first;
2848 out.assign(cols.size(), std::vector<std::size_t>());
2849 for (std::size_t j = 0; j < cols.size(); ++j) {
2850 for (std::size_t k = 0; k < cols[j].size(); ++k)
2851 for (std::size_t cc = 0; cc < m.chains.size(); ++cc)
2852 if (cols[j][k] < m.chains[cc].size() && m.chains[cc][cols[j][k]]) {
2853 bool seen = false;
2854 for (std::size_t q = 0; q < out[j].size(); ++q)
2855 if (out[j][q] == cc) seen = true;
2856 if (!seen) out[j].push_back(cc);
2857 }
2858 }
2859 }
2860 // A layer holding a POST_AND fork is a fork-join model like any other,
2861 // and the reference builds it as a `Network`, so buildLayers's
2862 // `getStruct()` runs the MMT node-visit pass on it: layer P:P2 of
2863 // lqn_workflows reports a Join visit of 4.3333, not the 1 the routing
2864 // solve leaves. THE LATER `refresh_chains()` CALLS DO NOT REPEAT IT, and
2865 // that is the reference's behaviour rather than an oversight here:
2866 // SolverLN.m re-runs `refreshChains()` alone after a routing reset,
2867 // which recomputes the visits WITHOUT the pass.
2869 }
2870
2871 /**
2872 * Port of layerSolverDeclaresRegion: true when the layer solver's feature
2873 * set covers an admission constraint. Asked while a layer is being built, so
2874 * it reads the solver CLASS's feature set, as the reference does: NC, FLD
2875 * and SSA declare Region, MVA does not.
2876 */
2877 bool layer_solver_declares_region() const {
2878 return opt.layer_solver == "nc" || opt.layer_solver == "fluid" || opt.layer_solver == "ssa";
2879 }
2880
2881 /**
2882 * Port of refPathAudit. Only the classes refpath wrote are checked: every
2883 * row they leave a node by must sum to one with no negative entry, and the
2884 * layer must carry exactly one reference class per chain it built, which is
2885 * what refresh_chains requires. Warnings, as in the reference.
2886 */
2887 template <class NeedsClass>
2888 void ref_path_audit(const qn::Layer<T>& m, std::size_t idx,
2889 const std::vector<std::size_t>& callers, const NeedsClass& needs_class,
2890 const std::vector<bool>& spliced, const std::vector<bool>& member,
2891 const std::vector<bool>& hop, const std::vector<bool>& expand,
2892 const std::vector<bool>& head_is_caller,
2893 const std::vector<std::size_t>& ref_stage,
2894 const std::vector<std::size_t>& hop_cls,
2895 const std::vector<std::size_t>& hop_ret,
2896 const std::vector<std::size_t>& ret_cls,
2897 const std::vector<std::size_t>& gate, const std::vector<std::size_t>& aux,
2898 const std::vector<std::size_t>& resume,
2899 const std::vector<std::size_t>& cls) const {
2900 std::set<std::size_t> watch;
2901 for (std::size_t k : ref_stage)
2902 if (k) watch.insert(k);
2903 for (std::size_t e = 1; e < hop.size(); ++e)
2904 if (hop[e]) {
2905 if (hop_cls[e]) watch.insert(hop_cls[e]);
2906 if (hop_ret[e]) watch.insert(hop_ret[e]);
2907 }
2908 for (std::size_t t = 1; t < member.size(); ++t)
2909 if (member[t]) {
2910 if (ret_cls[t]) watch.insert(ret_cls[t]);
2911 if (cls[t]) watch.insert(cls[t]);
2912 }
2913 for (std::size_t c = 1; c < expand.size(); ++c)
2914 if (expand[c]) {
2915 if (gate[c]) watch.insert(gate[c]);
2916 if (aux[c]) watch.insert(aux[c]);
2917 if (resume[c]) watch.insert(resume[c]);
2918 }
2919 const std::size_t I = m.nodes.size();
2920 for (std::size_t r : watch) {
2921 for (std::size_t n = 1; n <= I; ++n) {
2922 double tot = 0.0;
2923 bool neg = false;
2924 for (const auto& kv : m.P) {
2925 if (kv.first.first != r || n > kv.second.rows()) continue;
2926 for (std::size_t j = 0; j < kv.second.cols(); ++j) {
2927 const double v = dbl(kv.second(n - 1, j));
2928 if (v < -GlobalConstants::FineTol) neg = true;
2929 tot += v;
2930 }
2931 }
2932 if (neg)
2933 std::cerr << "[LINE] Warning: refpath layer '" << lqn.hashnames[idx]
2934 << "': class '" << m.classes[r - 1].name << "' routes out of node "
2935 << n << " with a negative probability.\n";
2936 if (tot > GlobalConstants::FineTol && std::fabs(tot - 1.0) > 1e-6)
2937 std::cerr << "[LINE] Warning: refpath layer '" << lqn.hashnames[idx]
2938 << "': class '" << m.classes[r - 1].name << "' leaves node " << n
2939 << " with total probability " << tot << ".\n";
2940 }
2941 }
2942 std::size_t nref = 0;
2943 for (const qn::JobClass& jc : m.classes)
2944 if (jc.is_ref_class) ++nref;
2945 std::size_t expect = 0;
2946 for (bool h : head_is_caller)
2947 if (!h) ++expect;
2948 for (std::size_t t : callers)
2949 if (!spliced[t] && needs_class(t)) ++expect;
2950 if (nref != expect)
2951 std::cerr << "[LINE] Warning: refpath layer '" << lqn.hashnames[idx] << "' declares "
2952 << nref << " reference classes where " << expect
2953 << " were built; a chain holding two makes refresh_chains throw.\n";
2954 }
2955
2956 /**
2957 * Port of routeSynchCall (buildLayersRecursive.m).
2958 *
2959 * A synchronous call is made `callmean` times per visit to the calling
2960 * activity, against `nreplicas` server replicas. The branch is chosen by
2961 * callmean ALONE, against 1 and never against nreplicas: a Bernoulli pass
2962 * carries at most one call, more than one needs the geometric loop through
2963 * the `<call>.Aux` class, and the replicas only SPLIT each probability by
2964 * ntgt -- they never relax that bound (buildLayersRecursive.m:1028-1051).
2965 * Comparing against nreplicas instead agrees with the reference only at
2966 * nreplicas == 1; on lqn_sockshop's replicated P2_1 it put a callmean of 1
2967 * on the < branch and doubled T1's throughput and P2_1's utilization.
2968 *
2969 * A second class is needed because a self-loop on the call class would also
2970 * re-enter its service. Which of the two carries the call time depends on
2971 * the direction:
2972 *
2973 * callmean < 1 the mean count is folded into the DEMAND
2974 * (callservt = callmean * W), so the call class must
2975 * be visited exactly once or the time is discounted
2976 * twice; .Aux absorbs the remaining branch.
2977 * callmean > 1 the call class is re-entered, and .Aux is the
2978 * return path that closes the loop with probability
2979 * 1/callmean, giving a mean of callmean visits.
2980 *
2981 * The four cases below are (job at client / at server) x (call targets an
2982 * entry of THIS layer's server / of some other task).
2983 */
2984 struct CtxPair {
2985 std::size_t curclass;
2986 int jobpos;
2987 };
2988
2989 /** The `<call>.Aux` suffix, which is how the reference identifies them too. */
2990 static bool is_aux_class(const std::string& name) {
2991 return name.size() >= 4 && name.compare(name.size() - 4, 4, ".Aux") == 0;
2992 }
2993 template <class Ctx>
2994 Ctx route_sync_call(qn::Layer<T>& m, std::size_t idx, std::size_t cidx, Ctx st,
2995 const std::vector<std::size_t>& tnode,
2996 const std::vector<std::size_t>& tstat,
2997 const std::vector<std::size_t>& callcls,
2998 const std::vector<std::size_t>& auxcallcls, std::size_t clientNode,
2999 int atClient, int atServer, bool flat) {
3000 const T one = Tone();
3001 // The callee's stations in THIS layer, empty when it is served
3002 // elsewhere: under `srvn` that is the layer's own server or nothing,
3003 // under `flat` it is whichever of the many servers the call targets.
3004 const std::size_t ntgt = tnode.size();
3005 const bool to_this_server = ntgt > 0;
3006 const T nrep = num_traits<T>::from_int(int(ntgt > 0 ? ntgt : 1));
3007 const T share = T(one / nrep);
3008 const T callmean = lqn.callproc_mean[cidx];
3009 const std::size_t cc = callcls[cidx];
3010 const std::size_t ax = auxcallcls[cidx];
3011 const bool below = callmean < one;
3012 const bool above = callmean > one;
3013 if (st.jobpos == atClient) {
3014 if (to_this_server) {
3015 if (below) {
3016 m.set_route(st.curclass, ax, clientNode, clientNode, T(one - callmean));
3017 for (std::size_t r = 0; r < ntgt; ++r) {
3018 m.set_route(st.curclass, cc, clientNode, tnode[r], T(callmean / nrep));
3019 m.set_route(cc, cc, tnode[r], clientNode, one);
3020 }
3021 // keeps .Aux attached to the routing graph; carries no time
3022 m.set_route(ax, cc, clientNode, clientNode, one);
3023 } else if (above) {
3024 for (std::size_t r = 0; r < ntgt; ++r) {
3025 m.set_route(st.curclass, cc, clientNode, tnode[r], share);
3026 m.set_route(cc, ax, tnode[r], clientNode, one);
3027 m.set_route(ax, cc, clientNode, tnode[r], T((one - one / callmean) / nrep));
3028 }
3029 m.set_route(ax, cc, clientNode, clientNode, T(one / callmean));
3030 } else {
3031 for (std::size_t r = 0; r < ntgt; ++r) {
3032 m.set_route(st.curclass, cc, clientNode, tnode[r], share);
3033 m.set_route(cc, cc, tnode[r], clientNode, one);
3034 }
3035 }
3036 for (std::size_t r = 0; r < ntgt; ++r) {
3037 m.set_service(tstat[r], cc, callservtproc[cidx]);
3038 call_map.push_back({idx, cidx, tstat[r], cc});
3039 }
3040 m.set_service(m.clientIdx, cc, Distrib<T>::immediate());
3041 st.jobpos = atClient;
3042 st.curnodes.clear();
3043 st.curclass = cc;
3044 } else {
3045 m.set_route(st.curclass, cc, clientNode, clientNode, one);
3046 if (below || above) {
3047 m.set_route(cc, ax, clientNode, clientNode, one);
3048 st.curclass = ax;
3049 } else {
3050 st.curclass = cc;
3051 }
3052 st.jobpos = atClient;
3053 st.curnodes.clear();
3054 m.set_service(m.clientIdx, cc, callservtproc[cidx]);
3055 call_map.push_back({idx, cidx, m.clientIdx, cc});
3056 }
3057 } else {
3058 // the node the job stands at, one per replica of wherever it landed
3059 auto fromNode = [&](std::size_t r) {
3060 return st.curnodes[std::min(r, st.curnodes.size() - 1)];
3061 };
3062 if (to_this_server) {
3063 if (below) {
3064 // The skip and the call merge back in the CALL class at the
3065 // client, exactly as in the atClient branch above. Routing
3066 // the skip into the call class instead and leaving in .Aux
3067 // gives .Aux no inbound arc at all, so its chain has no
3068 // reference class (buildLayersRecursive.m:1100-1118). This
3069 // port did that until 2026-08-11 and the arm is the same
3070 // under both layerings, as it is in the reference.
3071 for (std::size_t r = 0; r < ntgt; ++r) {
3072 m.set_route(st.curclass, ax, fromNode(r), clientNode, T(one - callmean));
3073 m.set_route(st.curclass, cc, fromNode(r), tnode[r], T(callmean / nrep));
3074 m.set_route(cc, cc, tnode[r], clientNode, one);
3075 }
3076 m.set_route(ax, cc, clientNode, clientNode, one);
3077 m.set_service(m.clientIdx, cc, Distrib<T>::immediate());
3078 st.jobpos = atClient;
3079 st.curnodes.clear();
3080 st.curclass = cc;
3081 } else if (above) {
3082 if (flat) {
3083 // the geometric repeat transits the client between
3084 // visits; a self-loop would merge them into one
3085 for (std::size_t r = 0; r < ntgt; ++r) {
3086 m.set_route(st.curclass, cc, fromNode(r), tnode[r], one);
3087 m.set_route(cc, ax, tnode[r], clientNode, one);
3088 m.set_route(ax, cc, clientNode, tnode[r],
3089 T((one - one / callmean) / nrep));
3090 }
3091 m.set_route(ax, cc, clientNode, clientNode, T(one / callmean));
3092 m.set_service(m.clientIdx, cc, Distrib<T>::immediate());
3093 st.curclass = cc;
3094 } else {
3095 for (std::size_t r = 0; r < ntgt; ++r) {
3096 m.set_route(st.curclass, cc, fromNode(r), tnode[r], one);
3097 m.set_route(cc, cc, tnode[r], tnode[r], T(one - one / callmean));
3098 m.set_route(cc, ax, tnode[r], clientNode, T(one / callmean));
3099 }
3100 st.curclass = ax;
3101 }
3102 st.jobpos = atClient;
3103 st.curnodes.clear();
3104 } else {
3105 for (std::size_t r = 0; r < ntgt; ++r)
3106 m.set_route(st.curclass, cc, fromNode(r), tnode[r], one);
3107 if (flat) {
3108 // the reply returns the job to the client, which is
3109 // where the successor restoration expects it
3110 for (std::size_t r = 0; r < ntgt; ++r)
3111 m.set_route(cc, cc, tnode[r], clientNode, one);
3112 m.set_service(m.clientIdx, cc, Distrib<T>::immediate());
3113 st.jobpos = atClient;
3114 st.curnodes.clear();
3115 } else {
3116 st.jobpos = atServer;
3117 st.curnodes = tnode;
3118 }
3119 st.curclass = cc;
3120 }
3121 for (std::size_t r = 0; r < ntgt; ++r) {
3122 m.set_service(tstat[r], cc, callservtproc[cidx]);
3123 call_map.push_back({idx, cidx, tstat[r], cc});
3124 }
3125 } else {
3126 for (std::size_t nd : st.curnodes)
3127 m.set_route(st.curclass, cc, nd, clientNode, one);
3128 if (below || above) {
3129 m.set_route(cc, ax, clientNode, clientNode, one);
3130 st.curclass = ax;
3131 } else {
3132 st.curclass = cc;
3133 }
3134 st.jobpos = atClient;
3135 st.curnodes.clear();
3136 m.set_service(m.clientIdx, cc, callservtproc[cidx]);
3137 call_map.push_back({idx, cidx, m.clientIdx, cc});
3138 }
3139 }
3140 return st;
3141 }
3142
3143 // -----------------------------------------------------------------------
3144 // getEntryServiceMatrix
3145 // -----------------------------------------------------------------------
3146
3147 /**
3148 * Reachability of activities and calls from each entry, as a 0/1 matrix over
3149 * the combined element-and-call index space. Multiplying it by
3150 * [residt; callresidt] sums an entry's whole service into one number.
3151 */
3152 void build_entry_service_matrix() {
3153 const std::size_t dim = lqn.nidx + lqn.ncalls;
3154 servtmatrix = Matrix<T>(dim + 1, dim + 1, Tzero());
3155 std::function<void(std::size_t, std::size_t)> rec = [&](std::size_t aidx, std::size_t eidx) {
3156 for (std::size_t nextaidx : lqn.graph.succ(aidx)) {
3157 const bool isLoop = lqn.graph.get(aidx, nextaidx) != lqn.dag.get(aidx, nextaidx);
3158 if (lqn.parent[aidx] != lqn.parent[nextaidx]) {
3159 for (std::size_t cidx : lqn.callsof[aidx])
3160 if (lqn.calltype[cidx] == CallType::SYNC)
3161 servtmatrix(eidx, lqn.nidx + cidx) = Tone();
3162 } else if (nextaidx != aidx && !isLoop) {
3163 servtmatrix(eidx, nextaidx) = Tone();
3164 rec(nextaidx, eidx);
3165 }
3166 }
3167 };
3168 for (std::size_t e = 1; e <= lqn.nentries; ++e) {
3169 const std::size_t eidx = lqn.eshift + e;
3170 rec(eidx, eidx);
3171 }
3172 }
3173
3174 // -----------------------------------------------------------------------
3175 // initInterlock
3176 // -----------------------------------------------------------------------
3177
3178 /** Port of initInterlock: the LQNS V5 static interlock analysis. */
3179 void init_interlock() {
3180 const std::size_t NE = lqn.nentries;
3181 il_all = Matrix<T>(NE + 1, NE + 1, Tzero());
3182 il_ph1 = Matrix<T>(NE + 1, NE + 1, Tzero());
3183
3184 std::function<void(std::size_t, std::size_t, T, T, std::vector<bool>&, int)> trace =
3185 [&](std::size_t eidx, std::size_t root_e, T pall, T pph1, std::vector<bool>& visited,
3186 int depth) {
3187 if (eidx <= lqn.eshift || eidx > lqn.eshift + NE) return;
3188 const std::size_t e = eidx - lqn.eshift;
3189 if (visited[e]) return;
3190 visited[e] = true;
3191 il_all(root_e, e) = T(il_all(root_e, e) + pall);
3192 il_ph1(root_e, e) = T(il_ph1(root_e, e) + pph1);
3193 for (std::size_t aidx : lqn.actsof[eidx]) {
3194 if (aidx <= lqn.ashift || aidx > lqn.ashift + lqn.nacts) continue;
3195 const std::size_t a = aidx - lqn.ashift;
3196 if (depth > 0 && lqn.actphase[a] > 1) continue;
3197 const bool is_ph1 = lqn.actphase[a] <= 1;
3198 for (std::size_t cidx : lqn.callsof[aidx]) {
3199 if (lqn.calltype[cidx] != CallType::SYNC) continue;
3200 if (!(lqn.callproc_mean[cidx] > Tzero())) continue;
3201 const std::size_t dst = lqn.callpair_dst[cidx];
3202 if (dst <= lqn.eshift || dst > lqn.eshift + NE) continue;
3203 trace(dst, root_e, T(pall * lqn.callproc_mean[cidx]),
3204 is_ph1 ? T(pph1 * lqn.callproc_mean[cidx]) : Tzero(), visited,
3205 depth + 1);
3206 }
3207 }
3208 visited[e] = false;
3209 };
3210 for (std::size_t e = 1; e <= NE; ++e) {
3211 std::vector<bool> visited(NE + 1, false);
3212 trace(lqn.eshift + e, e, Tone(), Tone(), visited, 0);
3213 }
3214
3215 il_common_entries.assign(NT() + 1, {});
3216 il_src_all.assign(NT() + 1, {});
3217 il_src_ph2.assign(NT() + 1, {});
3218 il_num_sources.assign(NT() + 1, 0.0);
3219
3220 for (std::size_t t = 1; t <= lqn.ntasks; ++t) {
3221 const std::size_t tidx = lqn.tshift + t;
3222 if (lqn.isref[tidx] || lqn.sched[tidx] == SchedStrategy::INF) continue;
3223 interlock_for_server(tidx);
3224 }
3225 for (std::size_t h = 1; h <= lqn.nhosts; ++h) {
3226 if (lqn.sched[h] == SchedStrategy::INF) continue;
3227 interlock_for_server(h);
3228 }
3229 }
3230
3231 std::vector<std::size_t> server_entry_nums(std::size_t serverIdx) const {
3232 std::vector<std::size_t> out;
3233 if (serverIdx <= lqn.nhosts) {
3234 for (std::size_t tidx : lqn.tasksof[serverIdx])
3235 for (std::size_t se : lqn.entriesof[tidx]) out.push_back(se - lqn.eshift);
3236 } else {
3237 for (std::size_t se : lqn.entriesof[serverIdx]) out.push_back(se - lqn.eshift);
3238 }
3239 return out;
3240 }
3241
3242 std::vector<std::size_t> client_tasks(std::size_t serverIdx) const {
3243 std::vector<std::size_t> out;
3244 if (serverIdx <= lqn.nhosts) return lqn.tasksof[serverIdx];
3245 for (std::size_t se : lqn.entriesof[serverIdx])
3246 for (std::size_t ci : lqn.iscaller.col(se))
3247 if (ci > lqn.tshift && ci <= NT()) out.push_back(ci);
3248 std::sort(out.begin(), out.end());
3249 out.erase(std::unique(out.begin(), out.end()), out.end());
3250 return out;
3251 }
3252
3253 std::vector<std::size_t> call_dst_tasks(std::size_t src_eidx, std::size_t target_e) const {
3254 std::vector<std::size_t> out;
3255 for (std::size_t aidx : lqn.actsof[src_eidx]) {
3256 if (aidx <= lqn.ashift || aidx > lqn.ashift + lqn.nacts) continue;
3257 for (std::size_t cidx : lqn.callsof[aidx]) {
3258 if (lqn.calltype[cidx] != CallType::SYNC) continue;
3259 const std::size_t dst = lqn.callpair_dst[cidx];
3260 const std::size_t de = dst - lqn.eshift;
3261 if (de >= 1 && de <= lqn.nentries && il_all(de, target_e) > Tzero())
3262 out.push_back(lqn.parent[dst]);
3263 }
3264 }
3265 std::sort(out.begin(), out.end());
3266 out.erase(std::unique(out.begin(), out.end()), out.end());
3267 return out;
3268 }
3269
3270 bool is_branch_point(std::size_t srcX, std::size_t entryA, std::size_t srcY,
3271 std::size_t entryB) const {
3272 const std::size_t taskA = lqn.parent[entryA], taskB = lqn.parent[entryB];
3273 const std::size_t taskX = lqn.parent[srcX];
3274 if (taskX == taskA && taskX == taskB) return false;
3275 if (srcX == entryA || srcY == entryB) return true;
3276 const std::vector<std::size_t> dx = call_dst_tasks(srcX, entryA - lqn.eshift);
3277 const std::vector<std::size_t> dy = call_dst_tasks(srcY, entryB - lqn.eshift);
3278 for (std::size_t a : dx)
3279 for (std::size_t b : dy)
3280 if (a != b) return true;
3281 return false;
3282 }
3283
3284 void trace_to_server(std::size_t eidx, std::size_t serverIdx, std::vector<bool>& visited,
3285 std::vector<std::size_t>& itasks, bool isHead) const {
3286 if (eidx <= lqn.eshift || eidx > lqn.eshift + lqn.nentries) return;
3287 const std::size_t e = eidx - lqn.eshift;
3288 if (visited[e]) return;
3289 const std::size_t owner = lqn.parent[eidx];
3290 if (owner == serverIdx) return;
3291 if (serverIdx <= lqn.nhosts && lqn.parent[owner] == serverIdx) return;
3292 visited[e] = true;
3293 bool found = false;
3294 const std::vector<std::size_t> sen = server_entry_nums(serverIdx);
3295 for (std::size_t aidx : lqn.actsof[eidx]) {
3296 if (aidx <= lqn.ashift || aidx > lqn.ashift + lqn.nacts) continue;
3297 for (std::size_t cidx : lqn.callsof[aidx]) {
3298 if (lqn.calltype[cidx] != CallType::SYNC) continue;
3299 const std::size_t dst = lqn.callpair_dst[cidx];
3300 const std::size_t dtask = lqn.parent[dst];
3301 bool reaches = dtask == serverIdx ||
3302 (serverIdx <= lqn.nhosts && lqn.parent[dtask] == serverIdx);
3303 if (!reaches) {
3304 const std::size_t de = dst - lqn.eshift;
3305 for (std::size_t sn : sen)
3306 if (il_all(de, sn) > Tzero()) reaches = true;
3307 }
3308 if (reaches) {
3309 trace_to_server(dst, serverIdx, visited, itasks, false);
3310 found = true;
3311 }
3312 }
3313 }
3314 if (found && !isHead) {
3315 itasks.push_back(owner);
3316 std::sort(itasks.begin(), itasks.end());
3317 itasks.erase(std::unique(itasks.begin(), itasks.end()), itasks.end());
3318 }
3319 visited[e] = false;
3320 }
3321
3322 void interlock_for_server(std::size_t serverIdx) {
3323 const std::vector<std::size_t> sen = server_entry_nums(serverIdx);
3324 if (sen.empty()) return;
3325 const std::vector<std::size_t> cts = client_tasks(serverIdx);
3326 if (cts.empty()) return;
3327
3328 std::vector<std::pair<std::size_t, std::size_t>> pairs; // (task, entry number)
3329 for (std::size_t ct : cts)
3330 for (std::size_t ce : lqn.entriesof[ct]) {
3331 const std::size_t cen = ce - lqn.eshift;
3332 if (cen < 1 || cen > lqn.nentries) continue;
3333 for (std::size_t se : sen)
3334 if (il_all(cen, se) > Tzero()) {
3335 pairs.emplace_back(ct, cen);
3336 break;
3337 }
3338 }
3339 if (pairs.size() < 2) return;
3340
3341 std::vector<std::size_t> common;
3342 for (std::size_t i = 0; i < pairs.size(); ++i)
3343 for (std::size_t j = i + 1; j < pairs.size(); ++j) {
3344 if (pairs[i].first == pairs[j].first) continue;
3345 const std::size_t eA = pairs[i].second, eC = pairs[j].second;
3346 for (std::size_t t = 1; t <= lqn.ntasks; ++t) {
3347 const std::size_t tidx = lqn.tshift + t;
3348 for (std::size_t ex : lqn.entriesof[tidx])
3349 for (std::size_t ey : lqn.entriesof[tidx]) {
3350 const std::size_t xn = ex - lqn.eshift, yn = ey - lqn.eshift;
3351 if (xn < 1 || yn < 1 || xn > lqn.nentries || yn > lqn.nentries) continue;
3352 if (il_all(xn, eA) > Tzero() && il_all(yn, eC) > Tzero() &&
3353 is_branch_point(ex, eA + lqn.eshift, ey, eC + lqn.eshift))
3354 common.push_back(ex);
3355 }
3356 }
3357 }
3358 std::sort(common.begin(), common.end());
3359 common.erase(std::unique(common.begin(), common.end()), common.end());
3360 if (common.empty()) return;
3361
3362 std::vector<std::size_t> interlocked;
3363 for (std::size_t ce : common) {
3364 std::vector<bool> visited(lqn.nentries + 1, false);
3365 std::vector<std::size_t> it;
3366 trace_to_server(ce, serverIdx, visited, it, true);
3367 for (std::size_t x : it) interlocked.push_back(x);
3368 }
3369 std::sort(interlocked.begin(), interlocked.end());
3370 interlocked.erase(std::unique(interlocked.begin(), interlocked.end()), interlocked.end());
3371
3372 std::vector<std::size_t> src_all;
3373 for (std::size_t ce : common) src_all.push_back(lqn.parent[ce]);
3374 std::sort(src_all.begin(), src_all.end());
3375 src_all.erase(std::unique(src_all.begin(), src_all.end()), src_all.end());
3376 {
3377 std::vector<std::size_t> diff;
3378 for (std::size_t x : src_all)
3379 if (!std::binary_search(interlocked.begin(), interlocked.end(), x)) diff.push_back(x);
3380 src_all = diff;
3381 }
3382
3383 // Left empty on purpose, and NOT because phase 2 is unsupported: the
3384 // reference computes ph2SrcTasks (initInterlock.m:230-262) and never
3385 // reads it, and its phase-2 source count is behind `if false`
3386 // (initInterlock.m:288-304). Filling it would change no number.
3387 std::vector<std::size_t> src_ph2;
3388 for (std::size_t it : interlocked)
3389 for (std::size_t ie : lqn.entriesof[it])
3390 for (std::size_t ci : lqn.iscaller.col(ie))
3391 if (ci > lqn.tshift && ci <= NT() &&
3392 !std::binary_search(interlocked.begin(), interlocked.end(), ci))
3393 src_all.push_back(ci);
3394 std::sort(src_all.begin(), src_all.end());
3395 src_all.erase(std::unique(src_all.begin(), src_all.end()), src_all.end());
3396
3397 double nsrc = 0.0;
3398 for (std::size_t st : src_all) nsrc += lqn.mult[st];
3399
3400 il_common_entries[serverIdx] = common;
3401 il_src_all[serverIdx] = src_all;
3402 il_src_ph2[serverIdx] = src_ph2;
3403 il_num_sources[serverIdx] = nsrc;
3404 }
3405
3406 // -----------------------------------------------------------------------
3407 // init and the outer iteration
3408 // -----------------------------------------------------------------------
3409 void init() {
3410 const std::size_t N = lqn.nidx;
3411 tput.assign(N + 1, Tzero());
3412 util.assign(N + 1, Tzero());
3413 servt.assign(N + 1, Tzero());
3414 residt.assign(N + 1, Tzero());
3415 thinkt.assign(N + 1, Tzero());
3416 callservt.assign(lqn.ncalls + 1, Tzero());
3417 callresidt.assign(lqn.ncalls + 1, Tzero());
3418 tputproc.assign(N + 1, Distrib<T>::disabled_dist());
3419 servt_ph1.assign(N + 1, Tzero());
3420 servt_ph2.assign(N + 1, Tzero());
3421 prOvertake.assign(lqn.nentries + 1, Tzero());
3422 build_entry_service_matrix();
3423
3424 relax_omega = (opt.relax == "fixed" || opt.relax == "adaptive") ? opt.relax_factor : 1.0;
3425
3426 servt_prev.assign(N + 1, std::numeric_limits<double>::quiet_NaN());
3427 residt_prev.assign(N + 1, std::numeric_limits<double>::quiet_NaN());
3428 tput_prev.assign(N + 1, std::numeric_limits<double>::quiet_NaN());
3429 thinkt_prev.assign(N + 1, std::numeric_limits<double>::quiet_NaN());
3430 callservt_prev.assign(lqn.ncalls + 1, std::numeric_limits<double>::quiet_NaN());
3431 callresidt_prev.assign(lqn.ncalls + 1, std::numeric_limits<double>::quiet_NaN());
3432 servt_prev_v.assign(N + 1, Tzero());
3433 residt_prev_v.assign(N + 1, Tzero());
3434 tput_prev_v.assign(N + 1, Tzero());
3435 thinkt_prev_v.assign(N + 1, Tzero());
3436 callservt_prev_v.assign(lqn.ncalls + 1, Tzero());
3437
3438 unique_route_idx.clear();
3439 for (const RouteRow& r : route_map) unique_route_idx.push_back(r.idx);
3440 std::sort(unique_route_idx.begin(), unique_route_idx.end());
3441 unique_route_idx.erase(std::unique(unique_route_idx.begin(), unique_route_idx.end()),
3442 unique_route_idx.end());
3443
3444 // only 'ilrate' reads the interlock tables, and their static analysis is costly
3445 if (interlockMethod == "ilrate") init_interlock();
3446
3447 maxitererr.assign(opt.iter_max + 2, 0.0);
3448 averagingstart = -1;
3449 hasconverged = false;
3450 moment_pass_done = false;
3451 results.clear();
3452
3453 servtcdf.assign(N + 1, fluid::FluidPassage());
3454 callservtcdf.assign(lqn.ncalls + 1, fluid::FluidPassage());
3455 entrycdfrespt.assign(lqn.nentries + 1, LnCdf());
3456 entryproc.assign(lqn.nentries + 1, mam::AphPair<T>());
3457 cdf_repo.assign(ensemble.size(), std::vector<std::vector<fluid::FluidPassage>>());
3458 }
3459
3460 /**
3461 * The Picard iteration, with the stochastic controller in place of the
3462 * deterministic test when the layer engine is a simulator.
3463 *
3464 * The two tests are NOT interchangeable and cannot both run: `converged`
3465 * folds a moving average into `results` in place, which is exactly what the
3466 * Polyak-Ruppert average would then be taken of a second time. Whichever
3467 * test is in force therefore owns the results.
3468 */
3469 void iterate() {
3470 init();
3471 const bool stoch = opt.layer_solver == "ssa";
3472 if (stoch) {
3473 LnStochConfig cfg;
3474 cfg.iter_tol = opt.iter_tol;
3475 cfg.relax_burnin = relax_omega;
3476 stoch_ctl = std::make_shared<LnStochController<T>>(cfg);
3477 }
3478 int it = 0;
3479 while (it < opt.iter_max) {
3480 if (!stoch && converged(it)) break;
3481 ++it;
3482 results.emplace_back(ensemble.size());
3483 for (std::size_t e = 0; e < ensemble.size(); ++e) analyze(it, e);
3484 post(it);
3485 if (stoch) {
3486 std::vector<double> jobs(ensemble.size(), 0.0);
3487 for (std::size_t e = 0; e < ensemble.size(); ++e)
3488 jobs[e] = ensemble[e].total_jobs();
3489 const bool stop = stoch_ctl->update(it, results.back(), jobs, servt, residt);
3490 // The step this sets is the one the NEXT updateMetrics applies,
3491 // which is why it is read back here and not before the sweep.
3492 relax_omega = stoch_ctl->relax_omega();
3493 if (stop) {
3494 did_converge = true;
3495 hasconverged = true;
3496 break;
3497 }
3498 }
3499 }
3500 iterations_done = it;
3501 // finish(): in Robbins-Monro mode report the averaged iterate, not the
3502 // last (noisy) one.
3503 if (stoch && stoch_ctl && stoch_ctl->averaging_count() > 0) {
3504 const std::vector<LayerResult<T>>& avg = stoch_ctl->averaged_results();
3505 for (std::size_t e = 0; e < results.back().size() && e < avg.size(); ++e)
3506 results.back()[e] = avg[e];
3507 servt = stoch_ctl->averaged_servt();
3508 residt = stoch_ctl->averaged_residt();
3509 }
3510 }
3511
3512 /**
3513 * Port of the filterMetric helper inside @@NetworkSolver/getAvg.m.
3514 *
3515 * This is not cosmetic post-processing: it is where a raw solver matrix
3516 * becomes the result the rest of LINE consumes, and three of its rules
3517 * change numbers that SolverLN then iterates on.
3518 *
3519 * 1. a (station, class) pair the class never visits is zeroed, whether
3520 * because its service is disabled or because its visit ratio is zero;
3521 * 2. anything below FineTol is snapped to zero;
3522 * 3. the caller's zeroMask is applied -- and for the queue length and the
3523 * utilization that mask is `RN < 10*FineTol`, which zeroes both
3524 * wherever the response time is immediate.
3525 *
3526 * Rule 3 is the one that matters most here. Every LQN entry, task and call
3527 * class is served by an Immediate distribution somewhere, whose response
3528 * time is 1e-8, so without it a layer reports the Immediate classes' share
3529 * of the population as real queue length and the LN think-time update reads
3530 * a task utilization that MATLAB reports as zero.
3531 */
3532 Matrix<T> filter_metric(const qn::Layer<T>& L, const Matrix<T>& metric,
3533 const std::vector<std::vector<bool>>* zero_mask) const {
3534 const std::size_t M = L.nstations, K = L.nclasses;
3535 Matrix<T> out(M, K, Tzero());
3536 for (std::size_t i = 0; i < M; ++i)
3537 for (std::size_t k = 0; k < K; ++k)
3538 if (!L.disabled[i][k]) out(i, k) = metric(i, k);
3539 if (zero_mask)
3540 for (std::size_t i = 0; i < M; ++i)
3541 for (std::size_t k = 0; k < K; ++k)
3542 if ((*zero_mask)[i][k]) out(i, k) = Tzero();
3543 for (std::size_t i = 0; i < M; ++i)
3544 for (std::size_t k = 0; k < K; ++k)
3545 if (dbl(out(i, k)) < GlobalConstants::FineTol) out(i, k) = Tzero();
3546 for (std::size_t k = 0; k < K; ++k) {
3547 std::size_t c = L.nchains;
3548 for (std::size_t cc = 0; cc < L.nchains; ++cc)
3549 if (L.chains[cc][k]) c = cc;
3550 if (c == L.nchains) continue;
3551 for (std::size_t i = 0; i < M; ++i)
3552 if (L.visits[c](L.stateful_of_station(i + 1) - 1, k) == Tzero())
3553 out(i, k) = Tzero();
3554 }
3555 return out;
3556 }
3557
3558 /**
3559 * Solve one layer, running the fork-join fixed point when it has a fork.
3560 *
3561 * The driver itself is shared (fj_driver.h); the layer path supplies
3562 * solver_mva_analyzer as the inner solve. The auxiliary arrival rates
3563 * fj_lambda are warm-started across outer iterations, per the reference.
3564 *
3565 * `default` becomes `amva` on a fork layer, which is what mvaDispatch.m does
3566 * for any model whose BASE has a fork: the transformed layer is mixed with
3567 * auxiliary near-zero-rate open classes, and the exact mixed MVA the default
3568 * ladder would otherwise pick degenerates on them.
3569 */
3570 /**
3571 * Run one layer through the fluid analyzer instead of MVA.
3572 *
3573 * WHAT MAKES THIS SOUND AT ALL. The layer handed to a solver is an ordinary
3574 * closed queueing network -- the layering has already replaced every call
3575 * by a class with a service demand -- so any NetworkSolver can solve it,
3576 * and the reference says exactly that by taking a solver factory. The outer
3577 * Picard iteration only ever reads [QN, UN, RN, TN] back, so the fluid
3578 * result maps onto the same shape MVA returns.
3579 *
3580 * WHAT CHANGES, and it is not nothing: the fluid limit is exact only as the
3581 * populations grow, so on a layer of a few jobs it is a genuinely different
3582 * approximation, not a slower route to the same fixed point. It also feeds
3583 * DIFFERENT service demands back into the next outer iteration, so the two
3584 * ensembles converge to different fixed points rather than to the same one
3585 * by different paths.
3586 *
3587 * REFUSED BY NAME: a non-double backend (LSODA is double-only, see
3588 * solver_fluid.h) and a fork layer (`fj_fixed_point` supplies MVA as its
3589 * inner solve, and the auxiliary near-zero-rate open classes it introduces
3590 * are not something the fluid drift represents).
3591 */
3592 /**
3593 * Solve one layer with CTMC and report it in the shape the MVA path returns.
3594 *
3595 * Only reached for a layer carrying a Region. `qn::Layer<T>` derives from
3596 * NetworkStruct, so the layer goes to the analyzer as it stands; the
3597 * generator needs the routing matrix, which the layer does not build until
3598 * asked (see solve_layer_fluid).
3599 */
3600 mva::MvaSolution<T> solve_layer_ctmc(std::size_t e) {
3601 qn::Layer<T>& L = ensemble[e];
3602 L.refresh_rt();
3603 // The capacities the WAITQ generator reads to bound each station's
3604 // marginal are filled by solve_layer, for every engine; the routing
3605 // matrix is not, because only this path and the fluid one need it.
3606 ctmc::CtmcOptions co;
3607 // _run_any, not _run: the plain generator implements DROP only, and a
3608 // region built from an LQN admission constraint carries the WAITQ
3609 // default, whose per-region token FIFO lives in the waitq generator.
3610 const mva::AvgResult<T> a = ctmc::solver_ctmc_run_analyzer_any(L, co);
3611 mva::MvaSolution<T> out;
3612 out.method = "ctmc";
3613 out.iter = 1;
3614 out.Q = a.QN;
3615 out.U = a.UN;
3616 out.R = a.RN;
3617 out.Tp = a.TN;
3618 out.C = a.CN;
3619 out.X = a.XN;
3620 return out;
3621 }
3622
3623 /**
3624 * Solve one layer with SolverMAM's `dec.poisson`, for a setup/delay-off.
3625 *
3626 * The method matters: `dec.poisson` caps the arrival superposition at one
3627 * phase AND skips the fork-join / retrial / ldqbd diversions, so the station
3628 * reaches the setup branch of solver_mam_basic rather than being taken by a
3629 * shape-matching special case. It is what SolverLN.m:183 asks for.
3630 */
3631 mva::MvaSolution<T> solve_layer_mam(std::size_t e) {
3632 qn::Layer<T>& L = ensemble[e];
3633 L.refresh_rt();
3634 L.refresh_capacity();
3635 mam::MamOptions mo;
3636 mo.method = "dec.poisson";
3637 mva::MvaSolution<T> out = mam::mam_dispatch(L, mo).sol;
3638 out.method = "mam";
3639 return out;
3640 }
3641
3642 /**
3643 * Solve one layer with SolverNC, the reference's `LN(model, @(m) NC(m))`.
3644 *
3645 * This is the layer engine the Java CLI's bare `ln` token selects, against
3646 * `ln.mva` for the MVA one. NC evaluates the normalizing constant rather
3647 * than the MVA recursion, so on a product-form layer it is the SAME answer
3648 * reached exactly instead of through the AMVA approximation -- and on a
3649 * layer that is not product form the two differ, which is the whole reason
3650 * the token exists.
3651 */
3652 mva::MvaSolution<T> solve_layer_nc(std::size_t e) {
3653 if (fj_tr[e].active())
3654 throw UnsupportedError(
3655 "SolverLN: layer '" + ensemble[e].name +
3656 "' carries a fork, whose fixed point is driven by MVA in this port; solve this "
3657 "model with layer_solver 'mva'");
3658 qn::Layer<T>& L = ensemble[e];
3659 L.refresh_rt();
3660 // `solver_nc_solve`, NOT `nc_dispatch`: the latter is the INNER network
3661 // solve, the one the caching-queueing decomposition itself calls as its
3662 // `netfun`, and it carries no cache branch at all. The three cache gates
3663 // live in `solver_nc_solve` (solver_nc_runner.h), which is the true
3664 // counterpart of `mva_dispatch` on the branch below. Routed here, a
3665 // cache layer never reached `solver_nc_cacheqn_analyzer` and its Cache
3666 // node was solved as ordinary routing, so the hit/miss split came back
3667 // as the probability `link()` offered -- an even 1/2, independent of
3668 // capacity, item count and replacement strategy, with no warning.
3669 // `.sol` alone is taken, exactly as the MVA branch takes it: the layer
3670 // reads the split off its own station throughputs, and the auxiliary
3671 // `refreshed_struct` both dispatchers also return is unused on either.
3672 mva::MvaSolution<T> out = nc::solver_nc_solve(L, opt.layer_nc).sol;
3673 out.method = "nc";
3674 return out;
3675 }
3676
3677 /**
3678 * Solve one layer by simulation, the reference's `LN(model, @(m) SSA(m))`.
3679 *
3680 * THE LAYER RESULTS ARE THEN NOISY, and the deterministic convergence test
3681 * cannot terminate against noise: its successive-difference error is
3682 * bounded below by the standard error of the estimates. `iterate` therefore
3683 * switches to `LnStochController` whenever this engine is selected, so the
3684 * relaxation decays as Robbins-Monro and the reported iterate is the
3685 * Polyak-Ruppert average. Selecting this engine and keeping the
3686 * deterministic test would simply run to iter_max.
3687 */
3688 mva::MvaSolution<T> solve_layer_ssa(std::size_t e) {
3689 if (fj_tr[e].active())
3690 throw UnsupportedError(
3691 "SolverLN: layer '" + ensemble[e].name +
3692 "' carries a fork, whose fixed point is driven by MVA in this port; solve this "
3693 "model with layer_solver 'mva'");
3694 qn::Layer<T>& L = ensemble[e];
3695 L.refresh_rt();
3696 const ssa::SsaSolution s = ssa::solver_ssa(L, opt.layer_ssa);
3697 mva::MvaSolution<T> out;
3698 out.method = "ssa";
3699 out.iter = 1;
3700 out.Q = Matrix<T>(L.nstations, L.nclasses, Tzero());
3701 out.U = Matrix<T>(L.nstations, L.nclasses, Tzero());
3702 out.R = Matrix<T>(L.nstations, L.nclasses, Tzero());
3703 out.Tp = Matrix<T>(L.nstations, L.nclasses, Tzero());
3704 for (std::size_t i = 0; i < L.nstations; ++i)
3705 for (std::size_t r = 0; r < L.nclasses; ++r) {
3706 out.Q(i, r) = num_traits<T>::from_double(s.QN(i, r));
3707 out.U(i, r) = num_traits<T>::from_double(s.UN(i, r));
3708 out.R(i, r) = num_traits<T>::from_double(s.RN(i, r));
3709 out.Tp(i, r) = num_traits<T>::from_double(s.TN(i, r));
3710 }
3711 out.C.assign(L.nclasses, Tzero());
3712 out.X.assign(L.nclasses, Tzero());
3713 for (std::size_t r = 0; r < L.nclasses && r < s.XN.size(); ++r) {
3714 out.X[r] = num_traits<T>::from_double(s.XN[r]);
3715 out.C[r] = num_traits<T>::from_double(s.CN[r]);
3716 }
3717 return out;
3718 }
3719
3720 mva::MvaSolution<T> solve_layer_fluid(std::size_t e) {
3721 if (fj_tr[e].active())
3722 throw UnsupportedError(
3723 "SolverLN: layer '" + std::to_string(e) +
3724 "' carries a fork, whose fixed point is driven by MVA in this port; solve this "
3725 "model with layer_solver 'mva'");
3726 qn::Layer<T>& L = ensemble[e];
3727 // THE LAYER HAS NO ROUTING MATRIX UNTIL IT IS ASKED FOR ONE. `buildLayers`
3728 // records the wiring with set_route and then calls refresh_chains, which
3729 // reads P directly and computes the VISITS; MVA needs nothing else, so
3730 // `rt` stays empty. Both fluid methods route through `sn.rt` instead, and
3731 // an empty one is not an error anywhere -- it is read as "no transition
3732 // exists", the drift decays to zero and every metric comes back 0. It is
3733 // rebuilt on every solve because update_routing_probabilities can change
3734 // the wiring between outer iterations.
3735 L.refresh_rt();
3736 mva::MvaSolution<T> out;
3737 out.method = "fluid";
3738 out.iter = 1;
3739 out.Q = Matrix<T>(L.nstations, L.nclasses, Tzero());
3740 out.U = Matrix<T>(L.nstations, L.nclasses, Tzero());
3741 out.R = Matrix<T>(L.nstations, L.nclasses, Tzero());
3742 out.Tp = Matrix<T>(L.nstations, L.nclasses, Tzero());
3743 out.C.assign(L.nclasses, Tzero());
3744 out.X.assign(L.nclasses, Tzero());
3745 detail::ln_fluid_solve(L, opt.layer_fluid, out);
3746 return out;
3747 }
3748
3749 mva::MvaSolution<T> solve_layer(std::size_t e) {
3750 // BEFORE ANY ENGINE, AND FOR EVERY ONE OF THEM. `build_layer` stops at
3751 // refresh_chains, so a layer reaches this point with `cap` and
3752 // `classcap` still EMPTY, while the reference hands SolverMVA a struct
3753 // that refreshCapacity has already filled. Every consumer of Kendall's
3754 // K -- `buffer_size`, and through it `has_blocking` and the BCMP gate
3755 // `has_product_form` -- indexes those vectors by station, so an empty
3756 // one is an out-of-range read and not a permissive default. It is
3757 // recomputed on each solve because the layer's class populations move
3758 // between outer iterations, and the buffers are derived from them.
3759 ensemble[e].refresh_capacity();
3760 // A layer carrying an admission constraint goes to CTMC whatever the
3761 // layer solver is, as adaptiveSolverFactory does (SolverLN.m:1047-1089):
3762 // MVA and the fluid path both refuse a Region by name, so leaving the
3763 // layer on them would refuse a model the reference solves. The
3764 // reference then falls back to LDES and SSA; neither is reachable here
3765 // (there is no cpp LDES, and cpp SSA refuses regions), so CTMC is the
3766 // only path and its state-space bound is this port's real limit.
3767 if (!ensemble[e].regions.empty()) return solve_layer_ctmc(e);
3768 // A setup no longer forces the MAM decomposition on the layer
3769 // (SolverLN.m, 2026-08-11): the cold start is charged to the entry by
3770 // setup_charge(), not wired into the station, so the layer is an
3771 // ordinary one and the user's own layer solver serves it.
3772 // A cache layer is dispatched to the integrated caching-queueing
3773 // analyzer, which reads the NODE routing to find what lies downstream
3774 // of the cache. The layer carries none until asked, and it has to be
3775 // rebuilt every pass because entry selection can move it. EVERY layer
3776 // engine below must reach that analyzer, not just the MVA one: this
3777 // refresh runs for all of them, and a branch that then hands the
3778 // rebuilt routing to a solver with no cache gate reports the offered
3779 // hit/miss split instead of the converged one.
3780 if (has_cache_node(e)) ensemble[e].refresh_rt();
3781 if (opt.layer_solver == "fluid") return solve_layer_fluid(e);
3782 if (opt.layer_solver == "nc") return solve_layer_nc(e);
3783 if (opt.layer_solver == "ssa") return solve_layer_ssa(e);
3784 if (opt.layer_solver != "mva")
3785 throw UnsupportedError("SolverLN: layer_solver '" + opt.layer_solver +
3786 "' is not available; use 'mva', 'nc', 'fluid' or 'ssa'");
3787 // Route each layer through the full mva_dispatch ladder, as MATLAB's
3788 // SolverLN does by handing every layer to SolverMVA.runAnalyzer: a plain
3789 // closed layer still lands on branch 13 (solver_mva_analyzer), but a
3790 // layer carrying load- or class-dependent scaling now reaches
3791 // solver_mvald_analyzer, matching the reference rather than silently
3792 // running the flat AMVA.
3793 mva::MvaOptions eopt = opt.layer;
3794 if (e < layer_interlock.size()) eopt.interlock = layer_interlock[e];
3795 if (!fj_tr[e].active())
3796 return mva::mva_dispatch(ensemble[e], eopt, layer_init_sol[e]).sol;
3797 mva::MvaOptions lopt = eopt;
3798 if (lopt.method == "default") lopt.method = "amva";
3799 return mva::fj_fixed_point(
3800 ensemble[e], fj_tr[e], fj_lambda[e], eopt,
3801 [&lopt](qn::NetworkStruct<T>& V) {
3802 return mva::mva_dispatch(V, lopt, Matrix<T>()).sol;
3803 });
3804 }
3805
3806 void analyze(int it, std::size_t e) {
3807 qn::Layer<T>& L = ensemble[e];
3808 const mva::MvaSolution<T> s = solve_layer(e);
3809 LayerResult<T>& r = results[it - 1][e];
3810 r.RN = filter_metric(L, s.R, nullptr);
3811 std::vector<std::vector<bool>> zmask(L.nstations, std::vector<bool>(L.nclasses, false));
3812 for (std::size_t i = 0; i < L.nstations; ++i)
3813 for (std::size_t k = 0; k < L.nclasses; ++k)
3814 zmask[i][k] = dbl(r.RN(i, k)) < 10.0 * GlobalConstants::FineTol;
3815 r.QN = filter_metric(L, s.Q, &zmask);
3816 r.UN = filter_metric(L, s.U, &zmask);
3817 r.TN = filter_metric(L, s.Tp, nullptr);
3818 r.WN = residence_from_response(L, r.RN);
3819 // warm start the next solve of this layer from the chain-aggregated
3820 // queue lengths, as the reference does for SolverMVA layers
3821 // A fork layer is exempt: the reference guards this on the result having
3822 // the LAYER's station count, and the transformed model reports one more
3823 // row (its Source), so the hint is never installed there. A fluid layer
3824 // is exempt too, and for the reference's own reason: `analyze` guards
3825 // the warm start on `strcmp(self.solvers{e}.name, 'SolverMVA')`, because
3826 // init_sol means a queue-length vector to AMVA and a full ODE state
3827 // vector to the fluid solver -- the two are not interchangeable.
3828 if (!fj_tr[e].active() && opt.layer_solver == "mva") {
3829 Matrix<T> Qch(L.nstations, L.nchains, Tzero());
3830 for (std::size_t c = 0; c < L.nchains; ++c)
3831 for (std::size_t i = 0; i < L.nstations; ++i) {
3832 T s2 = Tzero();
3833 for (std::size_t k : L.inchain[c]) s2 += s.Q(i, k - 1);
3834 Qch(i, c) = s2;
3835 }
3836 layer_init_sol[e] = Qch;
3837 }
3838 }
3839
3840 /** Port of sn_get_residt_from_respt: response time scaled by the visit ratio. */
3841 Matrix<T> residence_from_response(const qn::Layer<T>& L, const Matrix<T>& RN) const {
3842 Matrix<T> V(L.nstations, L.nclasses, Tzero());
3843 for (std::size_t c = 0; c < L.nchains; ++c)
3844 for (std::size_t i = 0; i < L.nstations; ++i) {
3845 const std::size_t sf = L.stateful_of_station(i + 1) - 1;
3846 for (std::size_t k = 0; k < L.nclasses; ++k)
3847 V(i, k) = T(V(i, k) + L.visits[c](sf, k));
3848 }
3849 Matrix<T> WN(L.nstations, L.nclasses, Tzero());
3850 for (std::size_t i = 0; i < L.nstations; ++i)
3851 for (std::size_t k = 0; k < L.nclasses; ++k) {
3852 if (L.disabled[i][k]) continue;
3853 if (!(RN(i, k) > Tzero())) continue;
3854 if (dbl(RN(i, k)) < GlobalConstants::FineTol) {
3855 WN(i, k) = RN(i, k);
3856 continue;
3857 }
3858 std::size_t c = 0;
3859 for (std::size_t cc = 0; cc < L.nchains; ++cc)
3860 if (L.chains[cc][k]) c = cc;
3861 const std::size_t rstat = L.classes[k].refstat;
3862 T den = Tzero();
3863 if (L.refclass[c] > 0) {
3864 den = V(rstat - 1, L.refclass[c] - 1);
3865 } else {
3866 for (std::size_t kk : L.inchain[c]) den += V(rstat - 1, kk - 1);
3867 }
3868 if (den == Tzero()) continue;
3869 WN(i, k) = T(RN(i, k) * V(i, k) / den);
3870 }
3871 for (std::size_t i = 0; i < L.nstations; ++i)
3872 for (std::size_t k = 0; k < L.nclasses; ++k)
3873 if (dbl(WN(i, k)) < 10.0 * GlobalConstants::FineTol) WN(i, k) = Tzero();
3874 return WN;
3875 }
3876
3877 void post(int it) {
3878 update_metrics(it);
3879 // Under 'refpath' the merged client chain performs the interlocking
3880 // structurally, and the discount as well would deflate the same wait twice
3881 if (interlockMethod == "ilrate") update_populations(it);
3882 update_think_times(it);
3883 update_layers(it);
3884 update_routing_probabilities(it);
3885 for (std::size_t e : route_reset) {
3886 ensemble[e].refresh_chains();
3887 layer_init_sol[e] = Matrix<T>();
3888 }
3889 // moment3 needs no refreshProcesses here: this port keeps the service
3890 // DISTRIBUTION on the layer and refresh_rates re-reads it, so there is no
3891 // second representation to fall out of step, unlike sn.proc in MATLAB.
3892 for (std::size_t e : svc_reset) ensemble[e].refresh_rates();
3893 }
3894
3895 /** Port of converged.m: moving average of the layer results plus the test. */
3896 bool converged(int it) {
3897 const std::size_t E = ensemble.size();
3898 const int iter_min = std::max<int>(2 * int(E), int(std::ceil(opt.iter_max / 4.0)));
3899 const int wnd_size = std::max(5, int(std::ceil(iter_min / 5.0)));
3900
3901 if (it >= iter_min && int(results.size()) >= wnd_size) {
3902 const T w = T(Tone() / num_traits<T>::from_int(wnd_size));
3903 for (std::size_t e = 0; e < E; ++e) {
3904 LayerResult<T>& cur = results[results.size() - 1][e];
3905 auto scale = [&](Matrix<T>& A) {
3906 for (std::size_t i = 0; i < A.rows(); ++i)
3907 for (std::size_t j = 0; j < A.cols(); ++j) A(i, j) = T(A(i, j) * w);
3908 };
3909 Matrix<T> Q = cur.QN, U = cur.UN, R = cur.RN, Tp = cur.TN, W = cur.WN;
3910 scale(Q); scale(U); scale(R); scale(Tp); scale(W);
3911 for (int k = 1; k < wnd_size; ++k) {
3912 const LayerResult<T>& old = results[results.size() - 1 - k][e];
3913 auto add = [&](Matrix<T>& A, const Matrix<T>& B) {
3914 for (std::size_t i = 0; i < A.rows(); ++i)
3915 for (std::size_t j = 0; j < A.cols(); ++j) A(i, j) = T(A(i, j) + B(i, j) * w);
3916 };
3917 add(Q, old.QN); add(U, old.UN); add(R, old.RN); add(Tp, old.TN); add(W, old.WN);
3918 }
3919 cur.QN = Q; cur.UN = U; cur.RN = R; cur.TN = Tp; cur.WN = W;
3920 }
3921 }
3922
3923 if (it > 1) {
3924 double err = 0.0;
3925 for (std::size_t e = 0; e < E; ++e) {
3926 const Matrix<T>& Q = results[results.size() - 1][e].QN;
3927 const Matrix<T>& Q1 = results[results.size() - 2][e].QN;
3928 const double Njobs = ensemble[e].total_jobs();
3929 if (!(Njobs > 0.0)) continue;
3930 double mx = 0.0;
3931 for (std::size_t i = 0; i < Q.rows(); ++i)
3932 for (std::size_t j = 0; j < Q.cols(); ++j)
3933 mx = std::max(mx, std::fabs(dbl(Q(i, j)) - dbl(Q1(i, j))));
3934 err += mx / Njobs;
3935 }
3936 maxitererr[it] = err;
3938 static_cast<long>(it),
3939 "layer iteration %zu: max queue-length change %.3e (tolerance %.3e)",
3940 static_cast<std::size_t>(it), err, opt.iter_tol);
3941 if (it == iter_min) {
3942 line::util::LineConsole::step("started averaging the iterates to aid convergence");
3943 averagingstart = it;
3944 }
3945 }
3946
3947 if (it > iter_min && maxitererr[it] < opt.iter_tol && maxitererr[it - 1] < opt.iter_tol &&
3948 maxitererr[it - 2] < opt.iter_tol) {
3949 if (!hasconverged) {
3950 hasconverged = true;
3951 } else {
3952 did_converge = true;
3953 return true;
3954 }
3955 } else {
3956 hasconverged = false;
3957 }
3958 return false;
3959 }
3960
3961 // -----------------------------------------------------------------------
3962 // updateMetrics
3963 // -----------------------------------------------------------------------
3964 /** Port of updateMetrics.m: the method selects which update runs. */
3965 // -----------------------------------------------------------------------
3966 // Method "srvn.ph": the activity graph of an entry as a phase-type server law
3967 //
3968 // Each layer is a two-station cycle, a client Delay plus the server, with
3969 // one closed class per caller task. The sequencing the default method
3970 // encodes as routing -- a class per entry, per activity and per call, plus
3971 // Fork, Join, Router and ClassSwitch nodes -- is composed instead into a
3972 // single phase-type service law per (layer, caller), by the exact
3973 // series-parallel reduction of Workflow. The layer therefore carries only
3974 // the client/server back-and-forth, and the activity graph survives as a
3975 // distribution.
3976 //
3977 // Port of the MATLAB @SolverLN/buildLayersPH.m, updateLayersPH.m,
3978 // updateMetricsPH.m, updateThinkTimesPH.m, phComposeEntryLaws.m and
3979 // getEnsembleAvgPH.m, of the JAR jline.solvers.ln.SolverLNPH and of the
3980 // Python line_solver.solvers.solver_ln.solver_ln_ph. See
3981 // _kb/06-solver-catalog.md (LN section) for the layering taxonomy.
3982 // -----------------------------------------------------------------------
3983
3984 /**
3985 * The method the layers were actually BUILT for: "srvn.ph", "srvn.cs",
3986 * "flat.cs", "flat.ph" or "moment3". Resolved once in build_layers, because
3987 * the alias "srvn" may fall back; every dispatch reads this and not
3988 * opt.method, so a reconstruction can never disagree with the layers it is
3989 * reading.
3990 */
3991 std::string lnmethod;
3992 /** True once ph_init_laws has composed the per-entry workflows. */
3993 bool ph_laws_ready = false;
3994
3995 /**
3996 * Normalise a method name onto one the solver dispatches on. A method name
3997 * carries TWO decisions: the LAYERING, which fixes what a submodel is, and
3998 * the ENCODING, which fixes how an activity graph is written into it.
3999 *
4000 * "srvn.cs" encodes the activity graph as ROUTING, "srvn.ph" as a composed
4001 * phase-type server law, "srvn" is the alias that takes "srvn.ph" where it
4002 * can serve the model and "srvn.cs" otherwise, "flat.cs" squashes every
4003 * server into one submodel with the routing encoding, "flat.ph" squashes
4004 * them with the composed one ("flat" is the alias of "flat.cs" and resolves
4005 * unconditionally rather than probing "flat.ph", because a model is squashed
4006 * in order to express what only the routing encoding carries), and "moment3"
4007 * is the three-moment distribution pass over the routing layers. "default"
4008 * is the srvn alias; an unrecognised method name takes "srvn.cs".
4009 */
4010 static std::string ln_requested_method(const std::string& method) {
4011 std::string m;
4012 for (char c : method) m += static_cast<char>(std::tolower(static_cast<unsigned char>(c)));
4013 if (m.empty() || m == "srvn" || m == "default" || m == "auto") return "srvn";
4014 if (m == "srvn.ph" || m == "ph") return "srvn.ph";
4015 if (m == "srvn.cs" || m == "srvncs" || m == "cs") return "srvn.cs";
4016 if (m == "flat.cs" || m == "flatcs" || m == "flat" || m == "squashed") return "flat.cs";
4017 if (m == "flat.ph" || m == "flatph" || m == "squashed.ph") return "flat.ph";
4018 if (m == "moment3") return "moment3";
4019 // An unrecognised token takes the routing encoding, which is what every
4020 // name other than "moment3" resolved to before the alias existed.
4021 return "srvn.cs";
4022 }
4023
4024 /** True when the layers are the collapsed phase-type ones of "srvn.ph". */
4025 bool is_srvn_ph() const { return lnmethod == "srvn.ph"; }
4026
4027 /**
4028 * True when the layers carry the COMPOSED phase-type server law rather than
4029 * the routing encoding of the activity graph, under either layering. The
4030 * encoding, not the layering, decides which update and reconstruction passes
4031 * run, so every such dispatch asks this and not for one method name.
4032 */
4033 bool is_ph_encoding() const { return lnmethod == "srvn.ph" || lnmethod == "flat.ph"; }
4034
4035 /**
4036 * Station of layer L that stands for LQN element ELEM, falling back to the
4037 * layer's own server when ELEM is not a server there. Under "srvn" the
4038 * fallback is the answer for every element; under "flat.cs" it is the map
4039 * that tells the many servers of one layer apart.
4040 */
4041 static std::size_t station_idx_of(const qn::Layer<T>& L, std::size_t elem) {
4042 if (elem >= 1 && elem < L.server_idx_of.size() && L.server_idx_of[elem] != 0)
4043 return L.server_idx_of[elem];
4044 return L.serverIdx;
4045 }
4046
4047 /**
4048 * Station of layer L that class K (0-based) is served at: the PROCESSOR of
4049 * an activity, the CALLED TASK of a call, the layer's own server otherwise.
4050 */
4051 std::size_t station_idx_of_class(const qn::Layer<T>& L, std::size_t k) const {
4052 const int kind = L.classes[k].attr_kind;
4053 const std::size_t a = L.classes[k].attr_idx;
4054 if (kind == int(LqnElement::ACTIVITY))
4055 return station_idx_of(L, lqn.parent[lqn.parent[a]]);
4056 if (kind == int(LqnElement::CALL))
4057 return station_idx_of(L, lqn.parent[lqn.callpair_dst[a]]);
4058 return L.serverIdx;
4059 }
4060
4061 /**
4062 * Answer whether "srvn.ph" can serve this model, without disturbing the
4063 * solver.
4064 *
4065 * Both the feature gate and the series-parallel reduction can refuse, and
4066 * the second only finds out by composing the per-entry workflows -- work the
4067 * build then reuses, since those laws do not depend on the iterate.
4068 */
4069 bool probe_srvn_ph() {
4070 try {
4071 ph_assert_supported();
4072 ph_init_laws();
4073 ph_laws_ready = true;
4074 return true;
4075 } catch (const std::exception&) {
4076 ph_laws_ready = false;
4077 return false;
4078 }
4079 }
4080
4081 /** One two-station layer, and the caller classes that cycle through it. */
4082 struct PHLayer {
4083 std::size_t idx = 0;
4084 bool ishost = false;
4085 std::vector<std::size_t> callers;
4086 /** 1-based class index of each caller task, 0 when absent. */
4087 std::vector<std::size_t> class_of_caller;
4088 std::size_t nreplicas = 1;
4089 /** 1-based station indices of the server replicas. */
4090 std::vector<std::size_t> qstations;
4091 /** Mean of the law each class is currently served with, by class index. */
4092 std::vector<T> svcmean_by_class;
4093 /** (class index, entry index) or (class index, -call index). */
4094 std::vector<std::pair<std::size_t, long>> open_arrivals;
4095 /**
4096 * Closed population of the MODEL this server sits in. Under "flat.ph"
4097 * that is every caller of the single network, not only the callers of
4098 * this one station, so it is recorded here rather than recomputed.
4099 */
4100 double npop = 0.0;
4101 };
4102
4103 // per-entry workflows and their composed laws
4104 std::vector<workflow::Workflow<T>> ph_wf, ph_wfhost;
4105 std::vector<std::unordered_map<std::size_t, T>> ph_execs;
4106 std::vector<bool> ph_has_wf;
4107 std::vector<workflow::PhLaw<T>> ph_hostlaw, ph_entrylaw;
4108 std::vector<T> ph_hostmean, ph_entrymean, ph_entryscv;
4109 std::vector<T> ph_share, ph_overlap, ph_setupshare, ph_xdemand;
4110 Matrix<T> ph_ncalls, ph_calltime;
4111 std::vector<T> ph_procresid, ph_actthinkt, ph_calltotal;
4112 std::vector<PHLayer> ph_layers; ///< by element index, empty where absent
4113 std::vector<bool> ph_has_layer;
4114
4115 /** Mean of an activity's own think time, 0 when it declares none. */
4116 T act_think_time(std::size_t aidx) const {
4117 if (aidx >= lqn.actthink.size() || lqn.actthink[aidx].disabled) return Tzero();
4118 const double m = dbl(lqn.actthink[aidx].mean);
4119 if (!std::isfinite(m) || m <= GlobalConstants::FineTol) return Tzero();
4120 return lqn.actthink[aidx].mean;
4121 }
4122
4123 /**
4124 * Mean cold start one request of task TIDX pays, 0 when it declares none.
4125 *
4126 * A SetupTask powers a thread down when it goes idle and pays a setup before
4127 * it can serve again. The thread is released at a reply and starts a
4128 * delay-off countdown D of mean d; it powers off only if D expires before
4129 * the next request arrives, and a request arriving first cancels the
4130 * countdown and pays nothing. With the idle interval I seen by one thread
4131 * and exponential D, p = P(D < I) = E[I]/(E[I]+d) and the charge is p*s.
4132 *
4133 * Admission takes an ACTIVE idle thread before it wakes a sleeping one, so
4134 * the pool that actually cycles is only as large as the load needs: with
4135 * offered load b = X*S = rho*mult threads, about max(1,b) stay hot, giving
4136 * E[I] = (max(1,b) - b)/X. Exact at mult = 1; above it the exact answer is
4137 * matrix-analytic (Gandhi, Harchol-Balter and Adan, Performance Evaluation
4138 * 67(11), 2010). Twin of MATLAB lqn_setup_charge.m, the JAR
4139 * SolverLN.setupCharge and the Python SolverLN._setup_charge.
4140 */
4141 double setup_charge(std::size_t tidx) const {
4142 if (tidx >= lqn.hassetup.size() || !lqn.hassetup[tidx]) return 0.0;
4143 const double s = tidx < lqn.setuptime.size() && !lqn.setuptime[tidx].disabled
4144 ? dbl(lqn.setuptime[tidx].mean) : 0.0;
4145 const double d = tidx < lqn.delayofftime.size() && !lqn.delayofftime[tidx].disabled
4146 ? dbl(lqn.delayofftime[tidx].mean) : 0.0;
4147 if (!(s > GlobalConstants::FineTol) || !(d > GlobalConstants::FineTol)) return 0.0;
4148 const double mult = lqn.mult[tidx];
4149 if (!std::isfinite(mult) || mult <= 0.0) return 0.0;
4150 if (tidx >= tput.size() || tidx >= util.size()) return s;
4151 const double X = dbl(tput[tidx]);
4152 if (!std::isfinite(X) || X <= GlobalConstants::FineTol) return s;
4153 double rho = dbl(util[tidx]);
4154 if (!std::isfinite(rho) || rho < 0.0) rho = 0.0;
4155 rho = std::min(rho, 1.0 - GlobalConstants::FineTol);
4156 const double b = rho * mult; // offered load, in threads
4157 const double EI = (std::max(1.0, b) - b) / X; // idle interval of a hot thread
4158 return s * EI / (EI + d);
4159 }
4160
4161 /**
4162 * Probability that a request for entry EIDX finds its task's thread off.
4163 * ONE closure for both methods: setup_charge returns p*s, so p is that over
4164 * s. It also answers p = 1 during construction, before the first solve has
4165 * sized tput or util.
4166 */
4167 double ph_setup_prob(std::size_t eidx) const {
4168 const std::size_t tidx = lqn.parent[eidx];
4169 if (tidx >= lqn.hassetup.size() || !lqn.hassetup[tidx]) return 0.0;
4170 const double s = tidx < lqn.setuptime.size() && !lqn.setuptime[tidx].disabled
4171 ? dbl(lqn.setuptime[tidx].mean) : 0.0;
4172 const double d = tidx < lqn.delayofftime.size() && !lqn.delayofftime[tidx].disabled
4173 ? dbl(lqn.delayofftime[tidx].mean) : 0.0;
4174 if (!(s > GlobalConstants::FineTol) || !(d > GlobalConstants::FineTol)) return 0.0;
4175 return std::min(1.0, std::max(0.0, setup_charge(tidx) / s));
4176 }
4177
4178 /** Divisor that scales a processor utilization into [0,1]. */
4179 double ph_host_servers(std::size_t hidx) const {
4180 if (lqn.sched[hidx] == SchedStrategy::INF) return 1.0;
4181 const double m = lqn.maxmult[hidx];
4182 return (std::isfinite(m) && m > 0.0) ? m : 1.0;
4183 }
4184
4185 /** True when any task calls entry EIDX, synchronously or not. */
4186 bool ph_any_caller_of(std::size_t eidx) const {
4187 return lqn.issynccaller.any_col(eidx) || lqn.isasynccaller.any_col(eidx);
4188 }
4189
4190 /**
4191 * True when an entry arrival is the ONLY way requests reach task TIDX.
4192 * "srvn.ph" refuses forwarding calls outright, so sync/async callers are the
4193 * whole test.
4194 */
4195 bool ph_open_arrival_only(std::size_t tidx) const {
4196 if (lqn.isref[tidx]) return false;
4197 for (std::size_t eidx : lqn.entriesof[tidx])
4198 if (ph_any_caller_of(eidx)) return false;
4199 for (std::size_t eidx : lqn.entriesof[tidx])
4200 if (lqn.has_arrival[eidx]) return true;
4201 return false;
4202 }
4203
4204 /** Asynchronous calls whose target entry belongs to TIDX. */
4205 std::vector<std::size_t> ph_async_calls_into(std::size_t tidx) const {
4206 std::vector<std::size_t> out;
4207 for (std::size_t cidx = 1; cidx <= lqn.ncalls; ++cidx) {
4208 if (lqn.calltype[cidx] != CallType::ASYNC) continue;
4209 for (std::size_t e : lqn.entriesof[tidx])
4210 if (lqn.callpair_dst[cidx] == e) { out.push_back(cidx); break; }
4211 }
4212 return out;
4213 }
4214
4215 /**
4216 * Features the collapsed layer cannot represent are refused by name rather
4217 * than silently degraded.
4218 */
4219 void ph_assert_supported(bool flat = false) const {
4220 // The list is a property of the ENCODING, so it is the same under either
4221 // layering; what the squashing adds on top is refused in ph_flat_server_set.
4222 const std::string mname = flat ? "flat.ph" : "srvn.ph";
4223 // PHASE 2 IS ASKED HERE AND NOWHERE ELSE. `has_phase2` is built during
4224 // layering rather than being a property of the model, so the report twin
4225 // below cannot ask it; every other rule is shared with it.
4226 if (has_phase2)
4227 throw UnsupportedError(
4228 "method='" + mname + "' does not support second-phase activities: the composed "
4229 "entry law has no reply point. Use method='default'.");
4230 // ONE RULE LIST, TWO CALLERS: this throws what `ph_method_refusal`
4231 // returns, in the same order, so the run and a report cannot say
4232 // different things about one model. The squashing rules it also carries
4233 // for `flat.ph` are the ones `ph_flat_server_set` raises, which runs
4234 // before this on that path, so the message a caller sees is unchanged.
4235 const std::string why = ph_method_refusal(mname);
4236 if (!why.empty()) throw UnsupportedError(why);
4237 }
4238
4239 /**
4240 * The same rules as a SENTENCE, so a report can withdraw a method it cannot run.
4241 *
4242 * ONE RULE LIST, TWO CALLERS. `ph_assert_supported` above answers the run
4243 * path by throwing; nothing answered the report, and `list_valid_methods`
4244 * returns the same eight names for every model, so every layered model was
4245 * offered every encoding and `srvn.ph`/`flat.ph` then threw on contact.
4246 *
4247 * PHASE 2 IS DELIBERATELY ABSENT. `has_phase2` is built during layering
4248 * rather than being a property of the model, so a gate cannot ask it without
4249 * doing the layering it is meant to precede; the run path still refuses it.
4250 *
4251 * @param method the concrete method name
4252 * @return empty string when the method can encode this model, else the reason
4253 */
4254 std::string ph_method_refusal(const std::string& method) const {
4255 if (method != "srvn.ph" && method != "flat.ph") return std::string();
4256 // The squashing refusals, `flat.ph` only: each carries PER-LAYER state
4257 // that one submodel cannot hold (ph_flat_server_set).
4258 if (method == "flat.ph") {
4259 for (std::size_t i = 1; i <= NT(); ++i) {
4260 if (lqn.repl[i] > 1.0)
4261 return "method='flat.ph' does not support replicated processors or tasks, "
4262 "whose replicas need a submodel each. Use method='srvn.ph'.";
4263 if (lqn.hassetup[i])
4264 return "method='flat.ph' does not support setup tasks, whose powered-down "
4265 "threads are per-layer state. Use method='srvn.ph'.";
4266 }
4267 }
4268 for (std::size_t cidx = 1; cidx <= lqn.ncalls; ++cidx)
4269 if (lqn.calltype[cidx] == CallType::FWD)
4270 return "method='" + method +
4271 "' does not support forwarding calls, whose target is not part of the "
4272 "caller's activity graph. Use method='default'.";
4273 for (std::size_t i = 1; i < lqn.iscache.size(); ++i)
4274 if (lqn.iscache[i])
4275 return "method='" + method +
4276 "' does not support cache tasks. Use method='default'.";
4277 for (std::size_t i = 1; i < lqn.hassetup.size(); ++i) {
4278 if (!lqn.hassetup[i]) continue;
4279 if (lqn.sched[i] == SchedStrategy::INF || !std::isfinite(lqn.mult[i]))
4280 return "method='" + method + "': task '" + lqn.names[i] +
4281 "' declares a setup time on an infinite-server task, which holds no "
4282 "thread to power down; give it a finite multiplicity.";
4283 }
4284 for (std::size_t i = 0; i < lqn.lincon_A.size(); ++i)
4285 if (lqn.lincon_A[i].rows() > 0)
4286 return "method='" + method +
4287 "' does not support admission constraints on a layer station. "
4288 "Use method='default'.";
4289 for (std::size_t i = 1; i < lqn.lldscaling.size(); ++i) {
4290 const char* fname = nullptr;
4291 if (!lqn.lldscaling[i].empty()) fname = "a load dependence";
4292 else if (lqn.cdscaling[i]) fname = "a class dependence";
4293 else if (lqn.jdscaling[i]) fname = "a joint dependence";
4294 else if (!lqn.pools[i].empty()) fname = "server pools";
4295 if (fname != nullptr)
4296 return "method='" + method +
4297 "' does not support queue-dependent service rates on a layer station ('" +
4298 lqn.names[i] + "' declares " + fname + "). Use method='srvn.cs'.";
4299 }
4300 return std::string();
4301 }
4302
4303 /** Build the per-entry workflows and the iteration-invariant processor law. */
4304 void ph_init_laws() {
4305 const std::size_t N = lqn.nidx;
4306 ph_wf.assign(N + 1, workflow::Workflow<T>("empty"));
4307 ph_wfhost.assign(N + 1, workflow::Workflow<T>("empty"));
4308 ph_execs.assign(N + 1, {});
4309 ph_has_wf.assign(N + 1, false);
4310 ph_hostlaw.assign(N + 1, workflow::PhLaw<T>());
4311 ph_entrylaw.assign(N + 1, workflow::PhLaw<T>());
4312 ph_hostmean.assign(N + 1, Tzero());
4313 ph_entrymean.assign(N + 1, Tzero());
4314 ph_entryscv.assign(N + 1, Tone());
4315 ph_share.assign(N + 1, Tzero());
4316 ph_overlap.assign(N + 1, Tone());
4317 ph_setupshare.assign(N + 1, Tzero());
4318 ph_xdemand.assign(N + 1, Tzero());
4319 ph_ncalls = Matrix<T>(N + 1, N + 1, Tzero());
4320 ph_calltime = Matrix<T>(N + 1, N + 1, Tzero());
4321 ph_procresid.assign(N + 1, Tzero());
4322 ph_actthinkt.assign(N + 1, Tzero());
4323 ph_calltotal.assign(N + 1, Tzero());
4324 ph_layers.assign(NT() + 1, PHLayer());
4325 ph_has_layer.assign(NT() + 1, false);
4326
4327 for (std::size_t e = 1; e <= lqn.nentries; ++e) {
4328 const std::size_t eidx = lqn.eshift + e;
4329 const std::size_t tidx = lqn.parent[eidx];
4330 if (ignore[tidx]) continue;
4331 api::lqn::EntryWorkflow<T> ew = api::lqn::entry_workflow(lqn, eidx, true);
4332 ph_wf[eidx] = std::move(ew.wf);
4333 ph_execs[eidx] = ew.execs;
4334 ph_has_wf[eidx] = true;
4335 api::lqn::EntryWorkflow<T> eh = api::lqn::entry_workflow(lqn, eidx, false);
4336 ph_wfhost[eidx] = std::move(eh.wf);
4337 // the processor sees the WORK of concurrent branches, not their
4338 // elapsed time, so the host law serialises an AND fork
4339 ph_hostlaw[eidx] = api::lqn::serial_law(ph_wfhost[eidx]);
4340 ph_hostmean[eidx] =
4341 api::lqn::ph_moments(ph_hostlaw[eidx].alpha, ph_hostlaw[eidx].S).first;
4342 }
4343
4344 // until the first iteration reports throughputs, a task splits its
4345 // requests evenly over its entries
4346 for (std::size_t t = 1; t <= lqn.ntasks; ++t) {
4347 const std::size_t tidx = lqn.tshift + t;
4348 const std::size_t n = lqn.entriesof[tidx].size();
4349 if (n == 0) continue;
4350 for (std::size_t eidx : lqn.entriesof[tidx])
4351 ph_share[eidx] = num_traits<T>::from_double(1.0 / double(n));
4352 }
4353 }
4354
4355 /**
4356 * Law of the total time one execution of the issuing activity spends in call
4357 * CIDX: the geometric compound, of mean callproc_mean, of the response law
4358 * of the called entry. The response law is fitted to the response time
4359 * reported by the callee's layer and to the SCV of the callee's own composed
4360 * law, so no extra solver output is needed.
4361 */
4362 Distrib<T> ph_call_burst_law(std::size_t cidx) const {
4363 const double m = dbl(lqn.callproc_mean[cidx]);
4364 const std::size_t eidx = lqn.callpair_dst[cidx];
4366 const T R = T(callservt[cidx] / lqn.callproc_mean[cidx]);
4367 double scvd = dbl(ph_entryscv[eidx]);
4368 if (!std::isfinite(scvd) || scvd <= GlobalConstants::FineTol) scvd = 1.0;
4369 const T Rf = dbl(R) > GlobalConstants::FineTol
4370 ? R : num_traits<T>::from_double(GlobalConstants::FineTol);
4371 const Distrib<T> base = lang::aph_fit_mean_scv(Rf, num_traits<T>::from_double(scvd));
4372 const workflow::PhLaw<T> body = api::lqn::ph_law_of(base);
4373 const workflow::PhLaw<T> loop =
4374 workflow::Workflow<T>::compose_loop_geometric(body, lqn.callproc_mean[cidx]);
4375 return Distrib<T>::phase_type(loop.alpha, loop.S,
4377 }
4378
4379 /**
4380 * Station law of a composed workflow. A geometric loop over a body of two or
4381 * more phases closes a cycle in the phase graph, and a cyclic generator is a
4382 * PH and not an APH: no layer solver declares PH, so such a law is reduced
4383 * to the APH with the SAME first two moments. AMVA and NC read exactly those
4384 * two, so the reduction is lossless for them and is a two-moment fit for the
4385 * phase-aware layer solvers.
4386 */
4387 Distrib<T> ph_station_law(const workflow::PhLaw<T>& law) const {
4389 return Distrib<T>::phase_type(law.alpha, law.S, true);
4390 const std::pair<T, T> mm = api::lqn::ph_moments(law.alpha, law.S);
4391 return lang::aph_fit_mean_scv(mm.first, mm.second);
4392 }
4393
4394 /** The law of a single immediate phase, the empty-composition answer. */
4395 static workflow::PhLaw<T> ph_immediate_law() {
4396 workflow::PhLaw<T> out;
4397 out.alpha.assign(1, Tone());
4398 out.S = Matrix<T>(1, 1, num_traits<T>::from_double(-GlobalConstants::Immediate));
4399 return out;
4400 }
4401
4402 /**
4403 * Recompose the entry service laws from the current fixed-point iterate.
4404 *
4405 * The composed mean is NOT the sum of the leaf means when the graph forks:
4406 * the branches of an AND fork overlap, and the entry finishes with the last
4407 * of them. The ratio of the two, the overlap factor, is what the caller-side
4408 * aggregates are scaled by, so that the pieces of a cycle still add up to
4409 * the cycle.
4410 */
4411 void ph_compose_entry_laws() {
4412 const std::size_t N = lqn.nidx;
4413 std::vector<T> entry_setup_share(N + 1, Tzero());
4414 ph_overlap.assign(N + 1, Tone());
4415
4416 for (std::size_t e = 1; e <= lqn.nentries; ++e) {
4417 const std::size_t eidx = lqn.eshift + e;
4418 if (!ph_has_wf[eidx]) continue;
4419 workflow::Workflow<T>& w = ph_wf[eidx];
4420 const std::unordered_map<std::size_t, T>& ex = ph_execs[eidx];
4421 T entrysum = Tzero(), procsum = Tzero();
4422 for (std::size_t aidx : lqn.actsof[eidx]) {
4423 T m = T(residt[aidx] + act_think_time(aidx));
4424 const T xa = ex.at(aidx);
4425 procsum += T(xa * m);
4426 const double md = dbl(m);
4427 w.set_activity_demand_mean(
4428 lqn.names[aidx],
4430 ? m : num_traits<T>::from_double(GlobalConstants::FineTol));
4431 for (std::size_t cidx : lqn.callsof[aidx]) {
4432 if (lqn.calltype[cidx] != CallType::SYNC) continue;
4433 w.set_activity_demand(lqn.callhashnames[cidx], ph_call_burst_law(cidx));
4434 m = T(m + callservt[cidx]);
4435 }
4436 entrysum += T(xa * m);
4437 }
4438 workflow::PhLaw<T> law = api::lqn::ph_law_of(w.refresh_ph());
4439 std::pair<T, T> mm = api::lqn::ph_moments(law.alpha, law.S);
4440 T m1 = mm.first, scv = mm.second;
4441 // All activities of an entry run on ONE processor, so the branches of
4442 // an AND fork cannot overlap the processor residence they request:
4443 // the composed maximum is a lower bound on the entry service time
4444 // only above that total. Where it falls below, the law is rescaled in
4445 // time to it, which keeps its shape, its SCV and its order.
4446 if (dbl(procsum) > dbl(m1) + GlobalConstants::FineTol) {
4447 const T f = T(m1 / procsum);
4448 for (std::size_t i = 0; i < law.S.rows(); ++i)
4449 for (std::size_t j = 0; j < law.S.cols(); ++j) law.S(i, j) = T(law.S(i, j) * f);
4450 m1 = procsum;
4451 }
4452 // A SetupTask powers a thread down when it goes idle, so a request may
4453 // find it off and pay a cold start before the entry runs at all. The
4454 // setup is not part of the activity graph and never enters the
4455 // series-parallel reduction: it is prefixed to the composed law
4456 // afterwards as the mixture p*(setup THEN entry) + (1-p)*entry, which
4457 // is again phase-type.
4458 const double p = ph_setup_prob(eidx);
4459 if (p > GlobalConstants::FineTol) {
4460 const std::size_t tidx = lqn.parent[eidx];
4461 const double sm = dbl(lqn.setuptime[tidx].mean);
4462 double sscv = dbl(lqn.setuptime[tidx].scv);
4463 if (!std::isfinite(sscv) || sscv <= GlobalConstants::FineTol) sscv = 1.0;
4464 if (std::isfinite(sm) && sm > GlobalConstants::FineTol) {
4465 const workflow::PhLaw<T> sl = api::lqn::ph_law_of(lang::aph_fit_mean_scv(
4466 num_traits<T>::from_double(sm), num_traits<T>::from_double(sscv)));
4467 std::vector<workflow::PhLaw<T>> mix;
4468 mix.push_back(workflow::Workflow<T>::compose_serial(sl, law));
4469 mix.push_back(law);
4470 std::vector<T> probs;
4471 probs.push_back(num_traits<T>::from_double(p));
4472 probs.push_back(num_traits<T>::from_double(1.0 - p));
4474 mm = api::lqn::ph_moments(law.alpha, law.S);
4475 m1 = mm.first;
4476 scv = mm.second;
4477 // The share of the entry law that is cold start and not work.
4478 // The surrogate-delay closure measures a thread's cycle in
4479 // WORK, so it must not read a station utilization this has
4480 // inflated -- see update_think_times_ph.
4481 const double denom = std::max(dbl(m1), GlobalConstants::FineTol);
4482 entry_setup_share[eidx] = num_traits<T>::from_double(p * sm / denom);
4483 }
4484 }
4485 ph_entrylaw[eidx] = law;
4486 ph_entrymean[eidx] = m1;
4487 ph_entryscv[eidx] = scv;
4488 if (dbl(entrysum) > GlobalConstants::FineTol) {
4489 const double r = std::min(1.0, dbl(m1) / dbl(entrysum));
4490 ph_overlap[eidx] = num_traits<T>::from_double(r);
4491 }
4492 }
4493
4494 // Per task, the share-weighted fraction of its station service that is
4495 // cold start rather than work.
4496 ph_setupshare.assign(N + 1, Tzero());
4497 for (std::size_t t = 1; t <= lqn.ntasks; ++t) {
4498 const std::size_t tidx = lqn.tshift + t;
4499 if (ignore[tidx]) continue;
4500 for (std::size_t eidx : lqn.entriesof[tidx])
4501 ph_setupshare[tidx] = T(ph_setupshare[tidx] + ph_share[eidx] * entry_setup_share[eidx]);
4502 }
4503
4504 // Expected number of calls per invocation, and the caller-side aggregates
4505 ph_ncalls = Matrix<T>(N + 1, N + 1, Tzero());
4506 ph_calltime = Matrix<T>(N + 1, N + 1, Tzero());
4507 ph_procresid.assign(N + 1, Tzero());
4508 ph_actthinkt.assign(N + 1, Tzero());
4509 ph_calltotal.assign(N + 1, Tzero());
4510 for (std::size_t t = 1; t <= lqn.ntasks; ++t) {
4511 const std::size_t tidx = lqn.tshift + t;
4512 if (ignore[tidx]) continue;
4513 for (std::size_t eidx : lqn.entriesof[tidx]) {
4514 if (!ph_has_wf[eidx]) continue;
4515 const T w = ph_share[eidx];
4516 if (!(dbl(w) > 0.0)) continue;
4517 const std::unordered_map<std::size_t, T>& ex = ph_execs[eidx];
4518 const T r = ph_overlap[eidx];
4519 for (std::size_t aidx : lqn.actsof[eidx]) {
4520 const T xa = ex.at(aidx);
4521 ph_procresid[tidx] = T(ph_procresid[tidx] + w * r * xa * residt[aidx]);
4522 ph_actthinkt[tidx] = T(ph_actthinkt[tidx] + w * r * xa * act_think_time(aidx));
4523 for (std::size_t cidx : lqn.callsof[aidx]) {
4524 if (lqn.calltype[cidx] != CallType::SYNC) continue;
4525 const std::size_t tgte = lqn.callpair_dst[cidx];
4526 const std::size_t tgtt = lqn.parent[tgte];
4527 // the COUNT of calls does not change with the overlap,
4528 // only the time the caller is held by them
4529 ph_ncalls(tidx, tgte) =
4530 T(ph_ncalls(tidx, tgte) + w * xa * lqn.callproc_mean[cidx]);
4531 ph_calltime(tidx, tgtt) =
4532 T(ph_calltime(tidx, tgtt) + w * r * xa * callservt[cidx]);
4533 ph_calltotal[tidx] = T(ph_calltotal[tidx] + w * r * xa * callservt[cidx]);
4534 }
4535 }
4536 }
4537 }
4538 }
4539
4540 /** Build the ensemble under method "srvn.ph". */
4541 void build_layers_ph(bool flat = false) {
4542 if (!ph_laws_ready) ph_assert_supported(flat);
4543 // The interlock correction rewrites the populations of the call classes,
4544 // which this method does not create: its callers reach the server in one
4545 // class each
4546 opt.interlocking = false;
4547 interlockMethod = "none";
4548
4549 // A preceding probe has already composed the per-entry workflows; they do
4550 // not depend on the iterate, so they are not rebuilt here.
4551 if (!ph_laws_ready) ph_init_laws();
4552
4553 // Seed the fixed point with the static demands, then compose the laws
4554 const std::size_t N = lqn.nidx;
4555 residt.assign(N + 1, Tzero());
4556 servt.assign(N + 1, Tzero());
4557 callservt.assign(lqn.ncalls + 1, Tzero());
4558 callresidt.assign(lqn.ncalls + 1, Tzero());
4559 for (std::size_t aidx = lqn.ashift + 1; aidx <= lqn.ashift + lqn.nacts; ++aidx)
4560 residt[aidx] = lqn.hostdem[aidx].disabled ? Tzero() : lqn.hostdem[aidx].mean;
4561 for (std::size_t cidx = 1; cidx <= lqn.ncalls; ++cidx) {
4562 if (lqn.calltype[cidx] != CallType::SYNC && lqn.calltype[cidx] != CallType::ASYNC)
4563 continue;
4564 const std::size_t eidx = lqn.callpair_dst[cidx];
4565 callservt[cidx] = T(lqn.callproc_mean[cidx] * ph_hostmean[eidx]);
4566 callresidt[cidx] = callservt[cidx];
4567 }
4568 ph_compose_entry_laws();
4569
4570 if (flat) {
4571 // ONE subnetwork holding every processor and every called task
4572 build_ph_flat_layer();
4573 tput.assign(N + 1, Tzero());
4574 util.assign(N + 1, Tzero());
4575 thinkt.assign(N + 1, Tzero());
4576 update_layers_ph(0);
4577 return;
4578 }
4579
4580 std::vector<qn::Layer<T>> raw(NT() + 1);
4581 std::vector<bool> present(NT() + 1, false);
4582
4583 for (std::size_t hidx = 1; hidx <= lqn.nhosts; ++hidx) {
4584 if (ignore[hidx]) continue;
4585 std::vector<std::size_t> callers;
4586 for (std::size_t tidx : lqn.tasksof[hidx]) {
4587 if (ignore[tidx]) continue;
4588 if (lqn.isref[tidx]) { callers.push_back(tidx); continue; }
4589 for (std::size_t eidx : lqn.entriesof[tidx])
4590 if (ph_any_caller_of(eidx) || lqn.has_arrival[eidx]) {
4591 callers.push_back(tidx);
4592 break;
4593 }
4594 }
4595 if (callers.empty()) continue;
4596 build_layer_ph(raw[hidx], hidx, callers, true);
4597 present[hidx] = true;
4598 }
4599 for (std::size_t t = 1; t <= lqn.ntasks; ++t) {
4600 const std::size_t tidx = lqn.tshift + t;
4601 if (ignore[tidx] || lqn.isref[tidx]) continue;
4602 std::vector<std::size_t> callers;
4603 for (std::size_t ct = 1; ct <= lqn.ntasks; ++ct) {
4604 const std::size_t c = lqn.tshift + ct;
4605 if (c == tidx || ignore[c]) continue;
4606 for (std::size_t e : lqn.entriesof[tidx])
4607 if (lqn.issynccaller.get(c, e)) { callers.push_back(c); break; }
4608 }
4609 if (callers.empty() && ph_async_calls_into(tidx).empty()) continue;
4610 build_layer_ph(raw[tidx], tidx, callers, false);
4611 present[tidx] = true;
4612 }
4613
4614 idxhash.assign(lqn.nidx + 1, -1);
4615 long next = 0;
4616 for (std::size_t i = 1; i <= NT(); ++i)
4617 if (present[i]) {
4618 idxhash[i] = next++;
4619 ensemble.push_back(std::move(raw[i]));
4620 }
4621 layer_init_sol.assign(ensemble.size(), Matrix<T>());
4622 build_fork_views();
4623
4624 // install the initial laws, so that iteration 1 sees the seeded demands
4625 // rather than the placeholders the stations were created with
4626 tput.assign(N + 1, Tzero());
4627 util.assign(N + 1, Tzero());
4628 thinkt.assign(N + 1, Tzero());
4629 update_layers_ph(0);
4630 }
4631
4632 /**
4633 * Build the ONE layer of method "flat.ph": a client delay plus a station for
4634 * every processor and every called task.
4635 *
4636 * A caller task is one closed class, and it visits each server it uses ONCE
4637 * per invocation, carrying there the composed law of the demand it places on
4638 * that server -- the same law method "srvn.ph" installs in the server's own
4639 * layer. What changes is that the servers now contend inside one network
4640 * instead of seeing each other through surrogate delays, so the client delay
4641 * keeps only the think times and whatever of the cycle this model does not
4642 * hold. That is the whole difference between the two encodings of the PH
4643 * composition, and it is why the reconstruction passes are shared verbatim.
4644 */
4645 void build_ph_flat_layer() {
4646 const std::vector<std::size_t> servers = ph_flat_server_set();
4647 const std::size_t nsrv = servers.size();
4648
4649 qn::Layer<T> m;
4650 m.name = "FlatPH";
4651 m.clientIdx = m.add_station(qn::Station<T>{
4652 "Clients", NodeType::Delay, SchedStrategy::INF,
4653 std::numeric_limits<double>::infinity(), false, 0});
4654 const std::size_t clientNode = m.node_of_station(m.clientIdx);
4655 m.server_idx_of.assign(lqn.nidx + 1, 0);
4656
4657 std::vector<std::size_t> station_of(lqn.nidx + 1, 0);
4658 std::vector<std::size_t> serverNode(nsrv, 0);
4659 std::vector<std::size_t> srvStation(nsrv, 0);
4660 for (std::size_t si = 0; si < nsrv; ++si) {
4661 const std::size_t idx = servers[si];
4662 const bool ishost = idx <= lqn.nhosts;
4663 qn::Station<T> st;
4664 st.name = lqn.hashnames[idx];
4665 st.nodetype = lqn.sched[idx] == SchedStrategy::INF ? NodeType::Delay : NodeType::Queue;
4666 st.sched = lqn.sched[idx];
4667 st.nservers = lqn.sched[idx] == SchedStrategy::INF
4668 ? std::numeric_limits<double>::infinity()
4669 : lqn.maxmult[idx];
4670 st.attr_ishost = ishost;
4671 st.attr_idx = idx;
4672 const std::size_t s = m.add_station(st);
4673 srvStation[si] = s;
4674 station_of[idx] = s;
4675 serverNode[si] = m.node_of_station(s);
4676 m.server_idx_of[idx] = s;
4677 if (ishost)
4678 m.host_stations.push_back(s);
4679 else
4680 m.task_stations.push_back(s);
4681 }
4682 // the scalar fallback of the station lookup, which no served element reaches
4683 m.serverIdx = station_of[servers[0]];
4684 m.flat = true;
4685
4686 // Callers of each server, and the union of them, which becomes the class set
4687 std::vector<std::vector<std::size_t>> callers_of(lqn.nidx + 1);
4688 std::vector<std::size_t> all_callers;
4689 for (std::size_t si = 0; si < nsrv; ++si) {
4690 const std::size_t idx = servers[si];
4691 std::vector<std::size_t> cs;
4692 if (idx <= lqn.nhosts) {
4693 for (std::size_t tidx : lqn.tasksof[idx]) {
4694 if (ignore[tidx]) continue;
4695 if (lqn.isref[tidx]) { cs.push_back(tidx); continue; }
4696 for (std::size_t eidx : lqn.entriesof[tidx])
4697 if (ph_any_caller_of(eidx) || lqn.has_arrival[eidx]) {
4698 cs.push_back(tidx);
4699 break;
4700 }
4701 }
4702 } else {
4703 for (std::size_t ct = 1; ct <= lqn.ntasks; ++ct) {
4704 const std::size_t c = lqn.tshift + ct;
4705 if (c == idx || ignore[c]) continue;
4706 for (std::size_t e : lqn.entriesof[idx])
4707 if (lqn.issynccaller.get(c, e)) { cs.push_back(c); break; }
4708 }
4709 }
4710 callers_of[idx] = cs;
4711 for (std::size_t c : cs)
4712 if (std::find(all_callers.begin(), all_callers.end(), c) == all_callers.end())
4713 all_callers.push_back(c);
4714 }
4715 std::sort(all_callers.begin(), all_callers.end());
4716
4717 // One closed class per caller task
4718 std::vector<std::size_t> class_of_caller(lqn.nidx + 1, 0);
4719 double npop = 0.0;
4720 for (std::size_t c : all_callers) {
4721 // ph_flat_server_set has refused every replicated element, so the
4722 // per-replica reduction the srvn builder makes is the identity here
4723 double nj = lqn.maxmult[c];
4724 if (std::isinf(nj)) {
4725 double sacc = 0.0;
4726 for (std::size_t k = 1; k <= NT(); ++k)
4727 if (lqn.taskgraph.get(k, c) != Tzero()) sacc += lqn.maxmult[k];
4728 nj = sacc;
4729 if (std::isinf(nj) || nj == 0.0) {
4730 double s2 = 0.0;
4731 for (std::size_t k = 1; k <= NT(); ++k)
4732 if (std::isfinite(lqn.maxmult[k])) s2 += lqn.maxmult[k];
4733 nj = std::min(s2, 1000.0);
4734 }
4735 }
4736 qn::JobClass jc;
4737 jc.name = lqn.hashnames[c];
4738 jc.type = JobClassType::CLOSED;
4739 jc.population = nj;
4740 jc.refstat = m.clientIdx;
4741 jc.is_ref_class = true;
4742 jc.attr_kind = int(LqnElement::TASK);
4743 jc.attr_idx = c;
4744 const std::size_t k = m.add_class(jc);
4745 class_of_caller[c] = k;
4746 m.attr_tasks.emplace_back(k, c);
4747 npop += nj;
4748 const double zt = dbl(ref_think_time(c));
4749 m.set_service(m.clientIdx, k,
4750 Distrib<T>::exp_mean(num_traits<T>::from_double(
4751 std::max(zt, GlobalConstants::FineTol))));
4752 // A station this caller never reaches must say so with a DISABLED law,
4753 // not with a tiny placeholder. An FCFS station carries ONE service law
4754 // across its classes, so a placeholder is not inert there: it is mixed
4755 // into the multiserver correction and invents waiting where there is
4756 // none. Under "srvn.ph" the question never arises, since every class of
4757 // a layer visits that layer's single server.
4758 for (std::size_t si = 0; si < nsrv; ++si)
4759 m.set_service(srvStation[si], k, Distrib<T>::disabled_dist());
4760 for (std::size_t si = 0; si < nsrv; ++si) {
4761 const std::size_t idx = servers[si];
4762 const std::vector<std::size_t>& cs = callers_of[idx];
4763 if (std::find(cs.begin(), cs.end(), c) == cs.end()) continue;
4764 m.set_service(srvStation[si], k,
4766 num_traits<T>::from_double(GlobalConstants::FineTol)));
4767 njobs(c, idx) = nj;
4768 thinkt_map.push_back({idx, c, m.clientIdx, k});
4769 servt_map.push_back({idx, c, station_of[idx], k});
4770 }
4771 }
4772
4773 // Open classes: entry arrivals on a processor station, async calls on a task one
4774 std::vector<std::vector<std::pair<std::size_t, long>>> open_of(lqn.nidx + 1);
4775 std::size_t sourceStation = 0, sinkNode = 0;
4776 for (std::size_t si = 0; si < nsrv; ++si) {
4777 const std::size_t hidx = servers[si];
4778 if (hidx > lqn.nhosts) continue;
4779 for (std::size_t c : callers_of[hidx]) {
4780 // A task no other task calls has no task station, so the think-time
4781 // closure never gives its caller class a surrogate delay: the class
4782 // cycles against an Immediate one and an open stream on top of it
4783 // doubles the load. The chain is the representation that honours the
4784 // thread pool, so it is kept and closed on the arrival rate instead.
4785 if (ph_open_arrival_only(c)) continue;
4786 for (std::size_t eidx : lqn.entriesof[c]) {
4787 if (!lqn.has_arrival[eidx]) continue;
4788 if (sourceStation == 0) {
4789 sourceStation = m.add_station(qn::Station<T>{
4790 "Source", NodeType::Source, SchedStrategy::EXT,
4791 std::numeric_limits<double>::infinity(), false, 0});
4792 m.sourceIdx = sourceStation;
4793 sinkNode = m.add_node("Sink", NodeType::Sink, false);
4794 m.sinkNode = sinkNode;
4795 }
4796 qn::JobClass oc;
4797 oc.name = lqn.hashnames[eidx] + ".Open";
4798 oc.type = JobClassType::OPEN;
4799 oc.population = std::numeric_limits<double>::infinity();
4800 oc.refstat = sourceStation;
4801 oc.attr_kind = int(LqnElement::ENTRY);
4802 oc.attr_idx = eidx;
4803 const std::size_t k = m.add_class(oc);
4804 m.set_service(sourceStation, k, lqn.arrival[eidx]);
4805 // disabled, not a placeholder, at every station this stream misses
4806 for (std::size_t s2 = 0; s2 < nsrv; ++s2)
4807 m.set_service(srvStation[s2], k, Distrib<T>::disabled_dist());
4808 const double hm = std::max(dbl(ph_hostmean[eidx]), GlobalConstants::FineTol);
4809 m.set_service(srvStation[si], k,
4810 Distrib<T>::exp_mean(num_traits<T>::from_double(hm)));
4811 open_of[hidx].emplace_back(k, long(eidx));
4812 m.attr_entries.emplace_back(k, eidx);
4813 }
4814 }
4815 }
4816 for (std::size_t si = 0; si < nsrv; ++si) {
4817 const std::size_t tidx = servers[si];
4818 if (tidx <= lqn.nhosts) continue;
4819 for (std::size_t cidx : ph_async_calls_into(tidx)) {
4820 if (sourceStation == 0) {
4821 sourceStation = m.add_station(qn::Station<T>{
4822 "Source", NodeType::Source, SchedStrategy::EXT,
4823 std::numeric_limits<double>::infinity(), false, 0});
4824 m.sourceIdx = sourceStation;
4825 sinkNode = m.add_node("Sink", NodeType::Sink, false);
4826 m.sinkNode = sinkNode;
4827 }
4828 const std::size_t eidx = lqn.callpair_dst[cidx];
4829 qn::JobClass oc;
4830 oc.name = lqn.callhashnames[cidx];
4831 oc.type = JobClassType::OPEN;
4832 oc.population = std::numeric_limits<double>::infinity();
4833 oc.refstat = sourceStation;
4834 oc.attr_kind = int(LqnElement::CALL);
4835 oc.attr_idx = cidx;
4836 const std::size_t k = m.add_class(oc);
4837 m.set_service(sourceStation, k, Distrib<T>::immediate());
4838 // disabled, not a placeholder, at every station this stream misses
4839 for (std::size_t s2 = 0; s2 < nsrv; ++s2)
4840 m.set_service(srvStation[s2], k, Distrib<T>::disabled_dist());
4841 const double em = std::max(dbl(ph_entrymean[eidx]), GlobalConstants::FineTol);
4842 m.set_service(srvStation[si], k,
4843 Distrib<T>::exp_mean(num_traits<T>::from_double(em)));
4844 open_of[tidx].emplace_back(k, -long(cidx));
4845 m.attr_calls.push_back({k, cidx, lqn.callpair_src[cidx], eidx});
4846 arv_call_map.push_back({tidx, cidx, sourceStation, k});
4847 call_map.push_back({tidx, cidx, station_of[tidx], k});
4848 }
4849 }
4850
4851 // Routing: one visit per server the caller uses, in server order. The
4852 // number of calls is carried by the service law, not by a visit ratio, so
4853 // no arc ever moves.
4854 for (std::size_t c : all_callers) {
4855 const std::size_t k = class_of_caller[c];
4856 std::size_t prev = clientNode;
4857 bool visited = false;
4858 for (std::size_t si = 0; si < nsrv; ++si) {
4859 const std::vector<std::size_t>& cs = callers_of[servers[si]];
4860 if (std::find(cs.begin(), cs.end(), c) == cs.end()) continue;
4861 m.set_route(k, k, prev, serverNode[si], Tone());
4862 prev = serverNode[si];
4863 visited = true;
4864 }
4865 if (visited) m.set_route(k, k, prev, clientNode, Tone());
4866 }
4867 if (sourceStation != 0) {
4868 const std::size_t srcNode = m.node_of_station(sourceStation);
4869 for (std::size_t si = 0; si < nsrv; ++si)
4870 for (const auto& oa : open_of[servers[si]]) {
4871 m.set_route(oa.first, oa.first, srcNode, serverNode[si], Tone());
4872 m.set_route(oa.first, oa.first, serverNode[si], sinkNode, Tone());
4873 }
4874 }
4875 apply_host_task_priorities(m, servers);
4876 m.refresh_chains();
4877
4878 idxhash.assign(lqn.nidx + 1, -1);
4879 for (std::size_t idx : servers) idxhash[idx] = 0;
4880 ensemble.clear();
4881 ensemble.push_back(std::move(m));
4882 layer_init_sol.assign(ensemble.size(), Matrix<T>());
4883 build_fork_views();
4884
4885 const std::size_t nclasses = ensemble[0].classes.size();
4886 for (std::size_t idx : servers) {
4887 PHLayer L;
4888 L.idx = idx;
4889 L.ishost = idx <= lqn.nhosts;
4890 L.callers = callers_of[idx];
4891 L.class_of_caller = class_of_caller;
4892 L.nreplicas = 1;
4893 L.qstations.assign(1, station_of[idx]);
4894 L.svcmean_by_class.assign(nclasses + 1, Tzero());
4895 L.open_arrivals = open_of[idx];
4896 L.npop = npop < 1.0 ? 1.0 : npop;
4897 ph_layers[idx] = L;
4898 ph_has_layer[idx] = true;
4899 }
4900 }
4901
4902 /**
4903 * Processors and called tasks that become stations of the flat layer.
4904 *
4905 * The set is the elements the srvn builder would have given a layer of their
4906 * own, so "flat.ph" and "srvn.ph" place the SAME stations and differ only in
4907 * how many networks hold them. The refusals are those of flat_server_set,
4908 * since they are properties of the squashing and not of the encoding: each of
4909 * these carries per-layer state that one submodel cannot hold.
4910 */
4911 std::vector<std::size_t> ph_flat_server_set() const {
4912 for (std::size_t i = 1; i <= NT(); ++i) {
4913 if (lqn.repl[i] > 1.0)
4914 throw UnsupportedError(
4915 "method='flat.ph' does not support replicated processors or tasks, whose "
4916 "replicas need a submodel each. Use method='srvn.ph'.");
4917 if (lqn.iscache[i])
4918 throw UnsupportedError(
4919 "method='flat.ph' does not support cache tasks. Use method='default'.");
4920 if (lqn.hassetup[i])
4921 throw UnsupportedError(
4922 "method='flat.ph' does not support setup tasks, whose powered-down threads "
4923 "are per-layer state. Use method='srvn.ph'.");
4924 }
4925 std::vector<std::size_t> servers;
4926 for (std::size_t hidx = 1; hidx <= lqn.nhosts; ++hidx) {
4927 if (ignore[hidx] || lqn.tasksof[hidx].empty()) continue;
4928 bool any = false;
4929 for (std::size_t tidx : lqn.tasksof[hidx]) {
4930 if (ignore[tidx]) continue;
4931 if (lqn.isref[tidx]) { any = true; break; }
4932 for (std::size_t eidx : lqn.entriesof[tidx])
4933 if (ph_any_caller_of(eidx) || lqn.has_arrival[eidx]) { any = true; break; }
4934 if (any) break;
4935 }
4936 if (any) servers.push_back(hidx);
4937 }
4938 for (std::size_t t = 1; t <= lqn.ntasks; ++t) {
4939 const std::size_t tidx = lqn.tshift + t;
4940 if (ignore[tidx] || lqn.isref[tidx]) continue;
4941 bool has_caller = false;
4942 for (std::size_t ct = 1; ct <= lqn.ntasks && !has_caller; ++ct) {
4943 const std::size_t c = lqn.tshift + ct;
4944 if (c == tidx || ignore[c]) continue;
4945 for (std::size_t e : lqn.entriesof[tidx])
4946 if (lqn.issynccaller.get(c, e)) { has_caller = true; break; }
4947 }
4948 if (!has_caller && ph_async_calls_into(tidx).empty()) continue;
4949 servers.push_back(tidx);
4950 }
4951 if (servers.empty())
4952 throw InputError(
4953 "method='flat.ph' found no server: the model has no processor with tasks.");
4954 return servers;
4955 }
4956
4957 /** Build the two-station layer of server element IDX. */
4958 void build_layer_ph(qn::Layer<T>& m, std::size_t idx,
4959 const std::vector<std::size_t>& callers, bool ishost) {
4960 m.name = lqn.hashnames[idx];
4961
4962 // Replicas of the server station, with the same fan-out reduction as the
4963 // default builder: a caller that reaches every replica sees one
4964 // representative
4965 const double rawrepl = lqn.repl[idx];
4966 std::size_t nreplicas = 1;
4967 if (rawrepl > 1.0 && !callers.empty()) {
4968 bool reduce = true;
4969 if (ishost) {
4970 for (std::size_t c : callers)
4971 if (lqn.repl[c] != rawrepl) reduce = false;
4972 } else {
4973 for (std::size_t c : callers)
4974 if (lqn.fanout_at(c, idx) < rawrepl) reduce = false;
4975 }
4976 nreplicas = reduce ? 1 : static_cast<std::size_t>(std::llround(rawrepl));
4977 if (reduce && !ishost) single_replica_tasks.insert(idx);
4978 }
4979 const bool reduce_fanout = (nreplicas == 1 && rawrepl > 1.0 && !callers.empty());
4980
4981 m.clientIdx = m.add_station(qn::Station<T>{
4982 "Clients", NodeType::Delay, SchedStrategy::INF,
4983 std::numeric_limits<double>::infinity(), false, 0});
4984 const std::size_t clientNode = m.node_of_station(m.clientIdx);
4985 m.serverIdx = m.clientIdx + 1;
4986 PHLayer L;
4987 L.idx = idx;
4988 L.ishost = ishost;
4989 L.callers = callers;
4990 L.nreplicas = nreplicas;
4991 L.class_of_caller.assign(lqn.nidx + 1, 0);
4992 std::vector<std::size_t> serverNode(nreplicas);
4993 for (std::size_t r = 0; r < nreplicas; ++r) {
4994 qn::Station<T> st;
4995 st.name = r == 0 ? lqn.hashnames[idx] : lqn.hashnames[idx] + "." + std::to_string(r + 1);
4996 st.nodetype = lqn.sched[idx] == SchedStrategy::INF ? NodeType::Delay : NodeType::Queue;
4997 st.sched = lqn.sched[idx];
4998 st.nservers = lqn.sched[idx] == SchedStrategy::INF
4999 ? std::numeric_limits<double>::infinity()
5000 : lqn.maxmult[idx];
5001 st.attr_ishost = ishost;
5002 st.attr_idx = idx;
5003 const std::size_t s = m.add_station(st);
5004 L.qstations.push_back(s);
5005 serverNode[r] = m.node_of_station(s);
5006 }
5007
5008 // --- closed class per caller task
5009 for (std::size_t c : callers) {
5010 double nj = njobs(c, idx);
5011 if (nj == 0.0) {
5012 const bool caller_single_replica =
5013 reduce_fanout || single_replica_tasks.count(c) > 0;
5014 nj = caller_single_replica ? lqn.maxmult[c] : lqn.maxmult[c] * lqn.repl[c];
5015 if (std::isinf(nj)) {
5016 double s = 0.0;
5017 for (std::size_t k = 1; k <= NT(); ++k)
5018 if (lqn.taskgraph.get(k, c) != Tzero()) s += lqn.maxmult[k];
5019 nj = s;
5020 if (std::isinf(nj) || nj == 0.0) {
5021 double s2 = 0.0;
5022 for (std::size_t k = 1; k <= NT(); ++k)
5023 if (std::isfinite(lqn.maxmult[k])) s2 += lqn.maxmult[k] * lqn.repl[k];
5024 nj = std::min(s2, 1000.0);
5025 }
5026 }
5027 njobs(c, idx) = nj;
5028 }
5029 qn::JobClass jc;
5030 jc.name = lqn.hashnames[c];
5031 jc.type = JobClassType::CLOSED;
5032 jc.population = nj;
5033 jc.refstat = m.clientIdx;
5034 jc.is_ref_class = true;
5035 jc.attr_kind = int(LqnElement::TASK);
5036 jc.attr_idx = c;
5037 const std::size_t k = m.add_class(jc);
5038 L.class_of_caller[c] = k;
5039 m.attr_tasks.emplace_back(k, c);
5040 const double zt = dbl(ref_think_time(c));
5041 m.set_service(m.clientIdx, k,
5042 Distrib<T>::exp_mean(num_traits<T>::from_double(
5043 std::max(zt, GlobalConstants::FineTol))));
5044 for (std::size_t s : L.qstations)
5045 m.set_service(s, k, Distrib<T>::exp_mean(
5046 num_traits<T>::from_double(GlobalConstants::FineTol)));
5047 // every layer must be refreshed after a law change: post() resets the
5048 // layers named by the think-time map
5049 thinkt_map.push_back({idx, c, m.clientIdx, k});
5050 servt_map.push_back({idx, c, L.qstations[0], k});
5051 }
5052
5053 // --- open classes: entry arrivals on a host layer, async calls on a task layer
5054 std::size_t sourceStation = 0, sinkNode = 0;
5055 if (ishost) {
5056 for (std::size_t c : callers) {
5057 // A task no other task calls has no task layer, so
5058 // update_think_times_ph never gives its caller class a surrogate
5059 // delay: the class cycles against an Immediate one and an open
5060 // stream on top of it doubles the load. The chain is the
5061 // representation that honours the thread pool, so it is kept and
5062 // closed on the arrival rate instead.
5063 if (ph_open_arrival_only(c)) continue;
5064 for (std::size_t eidx : lqn.entriesof[c]) {
5065 if (!lqn.has_arrival[eidx]) continue;
5066 if (sourceStation == 0) {
5067 sourceStation = m.add_station(qn::Station<T>{
5068 "Source", NodeType::Source, SchedStrategy::EXT,
5069 std::numeric_limits<double>::infinity(), false, 0});
5070 m.sourceIdx = sourceStation;
5071 sinkNode = m.add_node("Sink", NodeType::Sink, false);
5072 m.sinkNode = sinkNode;
5073 }
5074 qn::JobClass oc;
5075 oc.name = lqn.hashnames[eidx] + ".Open";
5076 oc.type = JobClassType::OPEN;
5077 oc.population = std::numeric_limits<double>::infinity();
5078 oc.refstat = sourceStation;
5079 oc.attr_kind = int(LqnElement::ENTRY);
5080 oc.attr_idx = eidx;
5081 const std::size_t k = m.add_class(oc);
5082 m.set_service(sourceStation, k, lqn.arrival[eidx]);
5083 for (std::size_t s : L.qstations) {
5084 const double hm = std::max(dbl(ph_hostmean[eidx]), GlobalConstants::FineTol);
5085 m.set_service(s, k, Distrib<T>::exp_mean(num_traits<T>::from_double(hm)));
5086 }
5087 L.open_arrivals.emplace_back(k, long(eidx));
5088 m.attr_entries.emplace_back(k, eidx);
5089 }
5090 }
5091 } else {
5092 for (std::size_t cidx : ph_async_calls_into(idx)) {
5093 if (sourceStation == 0) {
5094 sourceStation = m.add_station(qn::Station<T>{
5095 "Source", NodeType::Source, SchedStrategy::EXT,
5096 std::numeric_limits<double>::infinity(), false, 0});
5097 m.sourceIdx = sourceStation;
5098 sinkNode = m.add_node("Sink", NodeType::Sink, false);
5099 m.sinkNode = sinkNode;
5100 }
5101 const std::size_t eidx = lqn.callpair_dst[cidx];
5102 qn::JobClass oc;
5103 oc.name = lqn.callhashnames[cidx];
5104 oc.type = JobClassType::OPEN;
5105 oc.population = std::numeric_limits<double>::infinity();
5106 oc.refstat = sourceStation;
5107 oc.attr_kind = int(LqnElement::CALL);
5108 oc.attr_idx = cidx;
5109 const std::size_t k = m.add_class(oc);
5110 m.set_service(sourceStation, k, Distrib<T>::immediate());
5111 for (std::size_t s : L.qstations) {
5112 const double em = std::max(dbl(ph_entrymean[eidx]), GlobalConstants::FineTol);
5113 m.set_service(s, k, Distrib<T>::exp_mean(num_traits<T>::from_double(em)));
5114 }
5115 L.open_arrivals.emplace_back(k, -long(cidx));
5116 m.attr_calls.push_back({k, cidx, lqn.callpair_src[cidx], eidx});
5117 arv_call_map.push_back({idx, cidx, sourceStation, k});
5118 call_map.push_back({idx, cidx, L.qstations[0], k});
5119 }
5120 }
5121
5122 // Routing: one visit to the server per client cycle. The number of calls
5123 // is carried by the service law, not by a visit ratio, so no arc changes
5124 const T share = num_traits<T>::from_double(1.0 / double(nreplicas));
5125 for (std::size_t c : callers) {
5126 const std::size_t k = L.class_of_caller[c];
5127 for (std::size_t r = 0; r < nreplicas; ++r) {
5128 m.set_route(k, k, clientNode, serverNode[r], share);
5129 m.set_route(k, k, serverNode[r], clientNode, Tone());
5130 }
5131 }
5132 for (const auto& oa : L.open_arrivals) {
5133 const std::size_t k = oa.first;
5134 const std::size_t srcNode = m.node_of_station(sourceStation);
5135 for (std::size_t r = 0; r < nreplicas; ++r) {
5136 m.set_route(k, k, srcNode, serverNode[r], share);
5137 m.set_route(k, k, serverNode[r], sinkNode, Tone());
5138 }
5139 }
5140 L.svcmean_by_class.assign(m.classes.size() + 1, Tzero());
5141 double np = 0.0;
5142 for (std::size_t c : callers) {
5143 const double v = njobs(c, idx);
5144 if (std::isfinite(v) && v > 0.0) np += v;
5145 }
5146 L.npop = np < 1.0 ? 1.0 : np;
5147 apply_host_task_priorities(m, std::vector<std::size_t>{idx});
5148 m.refresh_chains();
5149 ph_layers[idx] = L;
5150 ph_has_layer[idx] = true;
5151 }
5152
5153 /** Law of the demand caller C places on the server of layer IDX per invocation. */
5154 workflow::PhLaw<T> ph_service_law(std::size_t idx, bool ishost, std::size_t c) const {
5155 if (ishost) {
5156 // mixture over the entries of C, weighted by their share of its requests
5157 std::vector<workflow::PhLaw<T>> laws;
5158 std::vector<T> probs;
5159 double tot = 0.0;
5160 for (std::size_t eidx : lqn.entriesof[c]) {
5161 if (ph_hostlaw[eidx].S.rows() == 0 || !(dbl(ph_share[eidx]) > 0.0)) continue;
5162 laws.push_back(ph_hostlaw[eidx]);
5163 probs.push_back(ph_share[eidx]);
5164 tot += dbl(ph_share[eidx]);
5165 }
5166 if (laws.empty()) return ph_immediate_law();
5167 for (T& p : probs) p = T(p / num_traits<T>::from_double(tot));
5168 return workflow::Workflow<T>::compose_mixture(laws, probs);
5169 }
5170 // task layer: the total demand is the sum, over the entries of the
5171 // server, of a geometric compound of the entry law of mean equal to the
5172 // number of calls
5173 bool started = false;
5174 workflow::PhLaw<T> out;
5175 for (std::size_t eidx : lqn.entriesof[idx]) {
5176 const T n = ph_ncalls(c, eidx);
5177 if (dbl(n) <= GlobalConstants::FineTol || ph_entrylaw[eidx].S.rows() == 0) continue;
5178 const workflow::PhLaw<T> lp =
5180 out = started ? workflow::Workflow<T>::compose_serial(out, lp) : lp;
5181 started = true;
5182 }
5183 if (!started) return ph_immediate_law();
5184 return out;
5185 }
5186
5187 /**
5188 * Mean time a thread of caller C spends away from the server of layer IDX per
5189 * invocation: idle, plus whatever of its cycle the layer does not hold.
5190 */
5191 T ph_delay_mean(std::size_t idx, std::size_t c) const {
5192 // ONE closure for both layerings. Under "srvn.ph" the model holds a single
5193 // server, so a host layer charges the whole call burst to the delay and a
5194 // task layer charges the caller's processor plus every other callee. Under
5195 // "flat.ph" the model holds every server, and only the think times are
5196 // left. Every term is SUMMED in rather than obtained by subtracting from a
5197 // total: that subtraction cancels catastrophically once a call time is
5198 // large, the think time falls below the ULP of the call time, and the layer
5199 // then sees a client delay of zero, saturates, and the fixed point runs
5200 // away.
5201 T z = T(thinkt[c] + ref_think_time(c));
5202 if (!std::isfinite(dbl(z)) || dbl(z) < 0.0) z = Tzero();
5203 z = T(z + ph_actthinkt[c]);
5204 const std::size_t hidx = lqn.parent[c];
5205 if (!ph_served_here(idx, hidx)) z = T(z + ph_procresid[c]);
5206 for (std::size_t t = 1; t <= lqn.ntasks; ++t) {
5207 const std::size_t tidx = lqn.tshift + t;
5208 if (!ph_served_here(idx, tidx)) z = T(z + ph_calltime(c, tidx));
5209 }
5210 if (!std::isfinite(dbl(z)) || dbl(z) < 0.0)
5211 z = num_traits<T>::from_double(GlobalConstants::FineTol);
5212 return z;
5213 }
5214
5215 /**
5216 * True when LQN element ELEM is a station of the same model that holds server
5217 * IDX. Under "srvn.ph" that is ELEM == IDX, since each server has a layer of
5218 * its own; under "flat.ph" it is every server of the one network.
5219 */
5220 bool ph_served_here(std::size_t idx, std::size_t elem) const {
5221 if (elem < 1 || elem >= idxhash.size() || idx < 1 || idx >= idxhash.size())
5222 return false;
5223 return idxhash[elem] >= 0 && idxhash[elem] == idxhash[idx];
5224 }
5225
5226 /**
5227 * Push the composed laws into the layers.
5228 *
5229 * A layer of this method carries no routing that depends on the iterate: the
5230 * number of calls a caller makes is folded into its service law rather than
5231 * into a visit ratio, so only two laws move per (layer, class) -- the
5232 * phase-type service law at the server and the mean of the surrogate delay
5233 * at the client.
5234 */
5235 void update_layers_ph(int) {
5236 for (std::size_t idx = 1; idx <= NT(); ++idx) {
5237 if (idxhash[idx] < 0 || !ph_has_layer[idx]) continue;
5238 PHLayer& L = ph_layers[idx];
5239 qn::Layer<T>& m = ensemble[std::size_t(idxhash[idx])];
5240 for (std::size_t c : L.callers) {
5241 const std::size_t k = L.class_of_caller[c];
5242 const workflow::PhLaw<T> sl = ph_service_law(idx, L.ishost, c);
5243 L.svcmean_by_class[k] = api::lqn::ph_moments(sl.alpha, sl.S).first;
5244 const Distrib<T> law = ph_station_law(sl);
5245 for (std::size_t s : L.qstations) m.set_service(s, k, law);
5246 const double zd = std::max(dbl(ph_delay_mean(idx, c)),
5248 m.set_service(m.clientIdx, k,
5249 Distrib<T>::exp_mean(num_traits<T>::from_double(zd)));
5250 }
5251 for (const auto& oa : L.open_arrivals) {
5252 const std::size_t k = oa.first;
5253 if (oa.second > 0) {
5254 // entry arrival: the processor demand law of the entry is static
5255 L.svcmean_by_class[k] = ph_hostmean[std::size_t(oa.second)];
5256 continue;
5257 }
5258 const std::size_t cidx = std::size_t(-oa.second);
5259 const std::size_t eidx = lqn.callpair_dst[cidx];
5260 L.svcmean_by_class[k] = ph_entrymean[eidx];
5261 const Distrib<T> law = ph_station_law(ph_entrylaw[eidx]);
5262 for (std::size_t s : L.qstations) m.set_service(s, k, law);
5263 const std::size_t aidx = lqn.callpair_src[cidx];
5264 double rate = dbl(tput[aidx]) * dbl(lqn.callproc_mean[cidx]);
5265 if (!std::isfinite(rate) || rate <= GlobalConstants::FineTol)
5267 m.set_service(m.sourceIdx, k,
5268 Distrib<T>::exp_rate(num_traits<T>::from_double(rate)));
5269 }
5270 m.refresh_rt();
5271 }
5272 }
5273
5274 /**
5275 * Residence time per visit, by Little from the queue length rather than from
5276 * the reported RN. A layer that saturates can come back from AMVA with an RN
5277 * no closed model can produce, and a reconstruction that trusts it feeds the
5278 * impossible value straight back into the call response times.
5279 */
5280 static T ph_residence(const T& Q, const T& X, const T& RN) {
5281 const double q = num_traits<T>::to_double(Q), x = num_traits<T>::to_double(X);
5282 if (std::isfinite(q) && q >= 0.0 && std::isfinite(x) && x > GlobalConstants::FineTol)
5283 return T(Q / X);
5284 return RN;
5285 }
5286
5287 /**
5288 * Ratio of a residence time to the mean of the law it was measured against,
5289 * bounded above by the layer population: a job can wait behind at most every
5290 * other job in a closed layer.
5291 */
5292 static T ph_inflation_of(const T& R, const T& S, double npop) {
5293 double f = 1.0;
5294 const double s = num_traits<T>::to_double(S), r = num_traits<T>::to_double(R);
5295 if (s > GlobalConstants::FineTol && std::isfinite(r) && r > 0.0) f = r / s;
5296 if (!std::isfinite(f) || f < 1.0) f = 1.0;
5297 if (std::isfinite(npop) && npop >= 1.0 && f > npop) f = npop;
5298 return num_traits<T>::from_double(f);
5299 }
5300
5301 /** Total closed population of a layer, i.e. how many jobs a job can queue behind. */
5302 double ph_layer_pop(const PHLayer& L, std::size_t idx) const {
5303 // Under "flat.ph" this is every caller of the single network and not only
5304 // the callers of this one station, so it is taken from the layer record.
5305 if (L.npop >= 1.0) return L.npop;
5306 double n = 0.0;
5307 for (std::size_t c : L.callers) {
5308 const double v = njobs(c, idx);
5309 if (std::isfinite(v) && v > 0.0) n += v;
5310 }
5311 return n < 1.0 ? 1.0 : n;
5312 }
5313
5314 /**
5315 * Reconstruct the LQN metrics.
5316 *
5317 * A layer of this method reports one row per caller task, not one per entry,
5318 * activity and call, so the per-element quantities the rest of SolverLN reads
5319 * -- servt, residt, callservt, callresidt, tput -- are recovered analytically
5320 * from the series-parallel weights of the entry workflows.
5321 *
5322 * The split is conservative by construction. A station reports a residence
5323 * time R per visit against a service law of mean S, so the queueing inflation
5324 * R/S is attributed to every leaf of that visit in proportion to its own
5325 * mean: the pieces sum back to R exactly.
5326 */
5327 void update_metrics_ph(int it) {
5328 const std::size_t N = lqn.nidx;
5329 servt.assign(N + 1, Tzero());
5330 residt.assign(N + 1, Tzero());
5331 callservt.assign(lqn.ncalls + 1, Tzero());
5332 callresidt.assign(lqn.ncalls + 1, Tzero());
5333
5334 std::vector<T> inflNum(N + 1, Tzero()), inflDen(N + 1, Tzero());
5335 std::vector<T> taskTput(N + 1, Tzero()), openTput(N + 1, Tzero());
5336
5337 // Host layers: the queueing inflation of the processor demand
5338 for (std::size_t hidx = 1; hidx <= lqn.nhosts; ++hidx) {
5339 if (idxhash[hidx] < 0 || !ph_has_layer[hidx]) continue;
5340 const PHLayer& L = ph_layers[hidx];
5341 const LayerResult<T>& res = results.back()[std::size_t(idxhash[hidx])];
5342 const double npop = ph_layer_pop(L, hidx);
5343 for (std::size_t c : L.callers) {
5344 const std::size_t k = L.class_of_caller[c];
5345 T X = Tzero(), Q = Tzero();
5346 for (std::size_t s : L.qstations) {
5347 X = T(X + res.TN(s - 1, k - 1));
5348 Q = T(Q + res.QN(s - 1, k - 1));
5349 }
5350 const T R = ph_residence(Q, X, res.RN(L.qstations[0] - 1, k - 1));
5351 const T f = ph_inflation_of(R, L.svcmean_by_class[k], npop);
5352 if (!std::isfinite(dbl(X)) || dbl(X) < 0.0) X = Tzero();
5353 // TOTAL over the replicas. The processor layer of a replicated
5354 // element models ONE representative replica, so X is one replica's
5355 // rate and the element's own rate is REPL times it. The matching
5356 // per-replica quantity is ph_xdemand, which the think-time closure
5357 // divides down for the same reason.
5358 taskTput[c] = T(taskTput[c] + num_traits<T>::from_double(lqn.repl[c]) * X);
5359 for (std::size_t eidx : lqn.entriesof[c]) {
5360 const T sh = dbl(ph_share[eidx]) > 0.0 ? ph_share[eidx] : Tzero();
5361 const T w = T(sh * X);
5362 inflNum[eidx] = T(inflNum[eidx] + w * f);
5363 inflDen[eidx] = T(inflDen[eidx] + w);
5364 }
5365 }
5366 for (const auto& oa : L.open_arrivals) {
5367 if (oa.second <= 0) continue; // an async call is served in the task layer
5368 const std::size_t k = oa.first;
5369 const std::size_t eidx = std::size_t(oa.second);
5370 T X = Tzero(), Q = Tzero();
5371 for (std::size_t s : L.qstations) {
5372 X = T(X + res.TN(s - 1, k - 1));
5373 Q = T(Q + res.QN(s - 1, k - 1));
5374 }
5375 if (!std::isfinite(dbl(X)) || dbl(X) <= 0.0) continue;
5376 const T f = ph_inflation_of(ph_residence(Q, X, res.RN(L.qstations[0] - 1, k - 1)),
5377 L.svcmean_by_class[k], npop);
5378 inflNum[eidx] = T(inflNum[eidx] + X * f);
5379 inflDen[eidx] = T(inflDen[eidx] + X);
5380 openTput[eidx] = T(openTput[eidx] + X);
5381 taskTput[lqn.parent[eidx]] = T(taskTput[lqn.parent[eidx]] + X);
5382 }
5383 }
5384
5385 for (std::size_t e = 1; e <= lqn.nentries; ++e) {
5386 const std::size_t eidx = lqn.eshift + e;
5387 double f = 1.0;
5388 if (dbl(inflDen[eidx]) > GlobalConstants::FineTol)
5389 f = dbl(inflNum[eidx]) / dbl(inflDen[eidx]);
5390 if (!std::isfinite(f) || f < 1.0) f = 1.0; // a residence cannot fall below its demand
5391 const T fv = num_traits<T>::from_double(f);
5392 for (std::size_t aidx : lqn.actsof[eidx])
5393 residt[aidx] = T(fv * (lqn.hostdem[aidx].disabled ? Tzero()
5394 : lqn.hostdem[aidx].mean));
5395 }
5396
5397 // Task layers: the response time of every call
5398 std::vector<T> relw(N + 1, Tzero());
5399 for (std::size_t t = 1; t <= lqn.ntasks; ++t) {
5400 const std::size_t tidx = lqn.tshift + t;
5401 if (idxhash[tidx] < 0 || !ph_has_layer[tidx]) continue;
5402 const PHLayer& L = ph_layers[tidx];
5403 const LayerResult<T>& res = results.back()[std::size_t(idxhash[tidx])];
5404 const double npop = ph_layer_pop(L, tidx);
5405 for (std::size_t c : L.callers) {
5406 const std::size_t k = L.class_of_caller[c];
5407 T X = Tzero(), Q = Tzero();
5408 for (std::size_t s : L.qstations) {
5409 X = T(X + res.TN(s - 1, k - 1));
5410 Q = T(Q + res.QN(s - 1, k - 1));
5411 }
5412 if (!std::isfinite(dbl(X)) || dbl(X) < 0.0) X = Tzero();
5413 const T g = ph_inflation_of(ph_residence(Q, X, res.RN(L.qstations[0] - 1, k - 1)),
5414 L.svcmean_by_class[k], npop);
5415 for (std::size_t cidx = 1; cidx <= lqn.ncalls; ++cidx) {
5416 if (lqn.calltype[cidx] != CallType::SYNC) continue;
5417 if (lqn.parent[lqn.callpair_src[cidx]] != c) continue;
5418 if (lqn.parent[lqn.callpair_dst[cidx]] != tidx) continue;
5419 const std::size_t eidx = lqn.callpair_dst[cidx];
5420 callservt[cidx] = T(lqn.callproc_mean[cidx] * g * ph_entrymean[eidx]);
5421 callresidt[cidx] = callservt[cidx];
5422 }
5423 for (std::size_t eidx : lqn.entriesof[tidx])
5424 relw[eidx] = T(relw[eidx] + X * ph_ncalls(c, eidx));
5425 }
5426 for (const auto& oa : L.open_arrivals) {
5427 if (oa.second >= 0) continue;
5428 const std::size_t k = oa.first;
5429 const std::size_t cidx = std::size_t(-oa.second);
5430 const std::size_t eidx = lqn.callpair_dst[cidx];
5431 T X = Tzero();
5432 for (std::size_t s : L.qstations) X = T(X + res.TN(s - 1, k - 1));
5433 const T R = res.RN(L.qstations[0] - 1, k - 1);
5434 if (std::isfinite(dbl(R)) && dbl(R) > 0.0) {
5435 callservt[cidx] = T(R * lqn.callproc_mean[cidx]);
5436 callresidt[cidx] = callservt[cidx];
5437 }
5438 if (std::isfinite(dbl(X)) && dbl(X) > 0.0) relw[eidx] = T(relw[eidx] + X);
5439 }
5440 }
5441
5442 // Entry shares and throughputs
5443 for (std::size_t t = 1; t <= lqn.ntasks; ++t) {
5444 const std::size_t tidx = lqn.tshift + t;
5445 const std::vector<std::size_t>& entries = lqn.entriesof[tidx];
5446 if (entries.empty()) continue;
5447 // How the requests SPLIT over the entries is a flow-balance question,
5448 // and is answered at the task layer: a caller class reaches that server
5449 // once per invocation of the caller, carrying its whole call burst in
5450 // its service law, so the station rate counts caller cycles and the
5451 // per-entry rate is that rate times the calls the caller makes.
5452 T tot = Tzero();
5453 for (std::size_t eidx : entries) tot = T(tot + relw[eidx] + openTput[eidx]);
5454 if (dbl(tot) > GlobalConstants::FineTol) {
5455 for (std::size_t eidx : entries)
5456 ph_share[eidx] = T((relw[eidx] + openTput[eidx]) / tot);
5457 } else {
5458 for (std::size_t eidx : entries)
5459 ph_share[eidx] = num_traits<T>::from_double(1.0 / double(entries.size()));
5460 }
5461 // HOW MANY requests the task completes is a different question, and the
5462 // flow-balance total does not answer it: that total is what the callers
5463 // DEMAND, not what the task's threads can deliver. A thread cycles
5464 // through its host demand AND then through the task think time, and only
5465 // the processor layer of the task carries both, so the rate is read there.
5466 tput[tidx] = dbl(taskTput[tidx]) > GlobalConstants::FineTol ? taskTput[tidx] : tot;
5467 for (std::size_t eidx : entries) tput[eidx] = T(tput[tidx] * ph_share[eidx]);
5468 // The DEMAND is kept apart because it, and not the rate just reported,
5469 // is what closes the surrogate delay: normalising the think time by a
5470 // rate the same think time produced makes the processor layer
5471 // self-referential. PER REPLICA, because the thread count it is paired
5472 // with there is per replica.
5473 const T nrep = num_traits<T>::from_double(std::max(1.0, lqn.repl[tidx]));
5474 ph_xdemand[tidx] = dbl(tot) > GlobalConstants::FineTol ? T(tot / nrep)
5475 : T(tput[tidx] / nrep);
5476 }
5477
5478 // Recovery, under-relaxation, and the derived per-element quantities
5479 for (std::size_t aidx = lqn.ashift + 1; aidx <= lqn.ashift + lqn.nacts; ++aidx) {
5480 T v = residt[aidx];
5481 if (!std::isfinite(dbl(v)) && it > 1 && !std::isnan(residt_prev[aidx]))
5482 v = residt_prev_v[aidx];
5483 if (relax_omega < 1.0 && it > 1 && !std::isnan(residt_prev[aidx])) {
5484 const T om = num_traits<T>::from_double(relax_omega);
5485 const T om1 = num_traits<T>::from_double(1.0 - relax_omega);
5486 v = T(om * v + om1 * residt_prev_v[aidx]);
5487 }
5488 residt[aidx] = v;
5489 residt_prev[aidx] = dbl(v);
5490 residt_prev_v[aidx] = v;
5491 }
5492 for (std::size_t cidx = 1; cidx <= lqn.ncalls; ++cidx) {
5493 T v = callservt[cidx];
5494 if (!std::isfinite(dbl(v)))
5495 v = (it > 1 && std::isfinite(callservt_prev[cidx])) ? callservt_prev_v[cidx]
5496 : Tzero();
5497 if (relax_omega < 1.0 && it > 1 && !std::isnan(callservt_prev[cidx])) {
5498 const T om = num_traits<T>::from_double(relax_omega);
5499 const T om1 = num_traits<T>::from_double(1.0 - relax_omega);
5500 v = T(om * v + om1 * callservt_prev_v[cidx]);
5501 }
5502 callservt[cidx] = v;
5503 callresidt[cidx] = v;
5504 callservt_prev[cidx] = dbl(v);
5505 callresidt_prev[cidx] = dbl(v);
5506 callservt_prev_v[cidx] = v;
5507 if (dbl(v) > 0.0) callservtproc[cidx] = Distrib<T>::exp_mean(v);
5508 }
5509
5510 // Recompose the entry laws from the iterate just computed. The entry
5511 // service time is then the mean of the COMPOSED law and not the sum of the
5512 // parts: the branches of an AND fork overlap, so an entry that forks
5513 // finishes with the last of its branches and is not charged their sum.
5514 ph_compose_entry_laws();
5515
5516 for (std::size_t e = 1; e <= lqn.nentries; ++e) {
5517 const std::size_t eidx = lqn.eshift + e;
5518 if (!ph_has_wf[eidx]) continue;
5519 const std::unordered_map<std::size_t, T>& ex = ph_execs[eidx];
5520 for (std::size_t aidx : lqn.actsof[eidx]) {
5521 T sa = T(residt[aidx] + act_think_time(aidx));
5522 for (std::size_t cidx : lqn.callsof[aidx])
5523 if (lqn.calltype[cidx] == CallType::SYNC) sa = T(sa + callservt[cidx]);
5524 servt[aidx] = sa;
5525 servt_prev[aidx] = dbl(sa);
5526 servt_prev_v[aidx] = sa;
5527 tput[aidx] = T(tput[eidx] * ex.at(aidx));
5528 tput_prev[aidx] = dbl(tput[aidx]);
5529 tput_prev_v[aidx] = tput[aidx];
5530 // exp_rate admits a null rate, but a never-called activity has no arrivals: disabled says so, as the python twin does.
5531 tputproc[aidx] = dbl(tput[aidx]) > 0.0 ? Distrib<T>::exp_rate(tput[aidx])
5532 : Distrib<T>::disabled_dist();
5533 if (dbl(sa) > 0.0) servtproc[aidx] = Distrib<T>::exp_mean(sa);
5534 }
5535 servt[eidx] = ph_entrymean[eidx];
5536 residt[eidx] = ph_entrymean[eidx];
5537 if (dbl(servt[eidx]) > 0.0) servtproc[eidx] = Distrib<T>::exp_mean(servt[eidx]);
5538 }
5539 }
5540
5541 /**
5542 * Surrogate delay of every caller.
5543 *
5544 * Same closure as update_think_times -- a thread of the task is idle for
5545 * whatever of its cycle the task's own station does not hold -- but the rate
5546 * it is normalised by is the INVOCATION rate of the task and not the
5547 * throughput of its station. Under this method a caller class reaches the
5548 * server once per invocation of the caller, carrying its whole call burst in
5549 * its service law, so the station rate counts caller cycles rather than calls
5550 * and the two differ by the mean number of calls.
5551 */
5552 void update_think_times_ph(int it) {
5553 thinktproc.assign(lqn.nidx + 1, Distrib<T>::disabled_dist());
5554 const T floorv = num_traits<T>::from_double(GlobalConstants::Zero);
5555 for (std::size_t t = 1; t <= lqn.ntasks; ++t) {
5556 const std::size_t tidx = lqn.tshift + t;
5557 if (ignore[tidx]) continue;
5558 const T ztask = ref_think_time(tidx);
5559 if (idxhash[tidx] < 0) {
5560 // A task no other task calls but whose entries carry an arrival
5561 // still has a cycle: its threads are driven by the stream.
5562 // build_layers_ph drops the open class for it precisely so this
5563 // closure can set the rate.
5564 const double arvrate = open_arrival_rate_of(tidx);
5565 if (arvrate > GlobalConstants::FineTol) {
5566 double nja = lqn.maxmult[tidx];
5567 if (!std::isfinite(nja) || nja <= 0.0) {
5568 nja = 0.0;
5569 for (std::size_t c = 1; c <= NT(); ++c) nja = std::max(nja, njobs(tidx, c));
5570 }
5571 T hres = Tzero();
5572 const std::size_t hidx = lqn.parent[tidx];
5573 if (hidx >= 1 && hidx < idxhash.size() && idxhash[hidx] >= 0 &&
5574 ph_has_layer[hidx] && ph_layers[hidx].class_of_caller[tidx] > 0) {
5575 const PHLayer& HL = ph_layers[hidx];
5576 const LayerResult<T>& hr = results.back()[std::size_t(idxhash[hidx])];
5577 const T rr = hr.RN(HL.qstations[0] - 1, HL.class_of_caller[tidx] - 1);
5578 if (!std::isnan(dbl(rr))) hres = rr;
5579 }
5580 T za = T(num_traits<T>::from_double(nja / arvrate) - hres - ztask);
5581 if (za < floorv) za = floorv;
5582 if (relax_omega < 1.0 && it > 1 && !std::isnan(thinkt_prev[tidx])) {
5583 const T om = num_traits<T>::from_double(relax_omega);
5584 const T om1 = num_traits<T>::from_double(1.0 - relax_omega);
5585 za = T(om * za + om1 * thinkt_prev_v[tidx]);
5586 }
5587 tput[tidx] = num_traits<T>::from_double(arvrate);
5588 thinkt[tidx] = za;
5589 thinkt_prev[tidx] = dbl(za);
5590 thinkt_prev_v[tidx] = za;
5591 thinktproc[tidx] = Distrib<T>::exp_mean(T(za + ztask));
5592 continue;
5593 }
5594 // a reference task, or one no other task calls: it has no station
5595 // of its own, so its only delay is the think time the user declared
5596 thinkt[tidx] = num_traits<T>::from_double(GlobalConstants::FineTol);
5597 thinktproc[tidx] = Distrib<T>::immediate();
5598 continue;
5599 }
5600 const PHLayer& L = ph_layers[tidx];
5601 const qn::Layer<T>& m = ensemble[std::size_t(idxhash[tidx])];
5602 const LayerResult<T>& r = results.back()[std::size_t(idxhash[tidx])];
5603 T U = Tzero();
5604 for (std::size_t k = 0; k < m.nclasses; ++k) {
5605 const T v = r.UN(L.qstations[0] - 1, k);
5606 if (!std::isnan(dbl(v))) U = T(U + v);
5607 }
5608 util[tidx] = U;
5609 // The closure below measures a thread's cycle in WORK: it is idle for
5610 // whatever of the cycle its station does not hold it working. A
5611 // SetupTask's station service also carries a cold start, which is time
5612 // the thread is unavailable but is not work, so it is taken back out of
5613 // U before the closure reads it. Zero for every task without a setup.
5614 if (dbl(ph_setupshare[tidx]) > 0.0)
5615 U = T(U * (Tone() - ph_setupshare[tidx]));
5616 // the rate the CALLERS ask of the task, not the rate its processor
5617 // layer reported: the latter is itself a function of this think time
5618 T X = ph_xdemand[tidx];
5619 if (!(dbl(X) > GlobalConstants::FineTol)) X = tput[tidx];
5620 // The thread pool of ONE replica, the convention ph_xdemand is kept in
5621 double nj = lqn.maxmult[tidx];
5622 if (!std::isfinite(nj) || nj <= 0.0) {
5623 nj = 0.0;
5624 for (std::size_t c = 1; c <= NT(); ++c) nj = std::max(nj, njobs(tidx, c));
5625 }
5626 T z;
5627 if (dbl(X) > GlobalConstants::FineTol) {
5628 if (lqn.sched[tidx] == SchedStrategy::INF) {
5629 // an infinite server reports a mean number of busy threads
5630 z = T((num_traits<T>::from_double(nj) - U) / X - ztask);
5631 } else {
5632 const T om = U > Tone() ? T(U - Tone()) : T(Tone() - U);
5633 z = T(num_traits<T>::from_double(nj) * om / X - ztask);
5634 }
5635 } else {
5636 z = thinkt[tidx];
5637 }
5638 if (z < floorv) z = floorv;
5639 if (it > 1 && !std::isnan(thinkt_prev[tidx]) && !std::isfinite(dbl(z)))
5640 z = thinkt_prev_v[tidx];
5641 if (relax_omega < 1.0 && it > 1 && !std::isnan(thinkt_prev[tidx])) {
5642 const T om = num_traits<T>::from_double(relax_omega);
5643 const T om1 = num_traits<T>::from_double(1.0 - relax_omega);
5644 z = T(om * z + om1 * thinkt_prev_v[tidx]);
5645 }
5646 thinkt[tidx] = z;
5647 thinkt_prev[tidx] = dbl(z);
5648 thinkt_prev_v[tidx] = z;
5649 thinktproc[tidx] = Distrib<T>::exp_mean(T(z + ztask));
5650 }
5651 }
5652
5653 /**
5654 * LQN-level results of method "srvn.ph".
5655 *
5656 * The layers report per caller task, so every entry, activity and call figure
5657 * is rebuilt from the converged fixed point rather than read off a class row,
5658 * in the layout aggregate() returns: QN carries the entry and task
5659 * utilizations, UN the processor utilizations, RN the response times and WN
5660 * the residence times.
5661 */
5662 LnSolution<T> aggregate_ph() {
5663 const std::size_t N = lqn.nidx;
5664 LnSolution<T> out;
5665 out.QN.assign(N + 1, Tzero());
5666 out.UN.assign(N + 1, Tzero());
5667 out.RN.assign(N + 1, Tzero());
5668 out.TN.assign(N + 1, Tzero());
5669 out.AN.assign(N + 1, Tzero());
5670 out.WN.assign(N + 1, Tzero());
5671 out.defined_Q.assign(N + 1, false);
5672 out.defined_U.assign(N + 1, false);
5673 out.defined_R.assign(N + 1, false);
5674 out.defined_T.assign(N + 1, false);
5675 out.defined_A.assign(N + 1, false);
5676 out.defined_W.assign(N + 1, false);
5677 out.iterations = iterations_done;
5678 out.converged = did_converge;
5679
5680 std::vector<T> PN(N + 1, Tzero()), UT(N + 1, Tzero());
5681 std::vector<bool> hasPN(N + 1, false), hasUT(N + 1, false);
5682
5683 for (std::size_t a = 1; a <= lqn.nacts; ++a) {
5684 const std::size_t aidx = lqn.ashift + a;
5685 const std::size_t tidx = lqn.parent[aidx];
5686 if (ignore[tidx]) continue;
5687 const std::size_t hidx = lqn.parent[tidx];
5688 out.TN[aidx] = tput[aidx];
5689 out.defined_T[aidx] = true;
5690 out.RN[aidx] = servt[aidx];
5691 out.defined_R[aidx] = true;
5692 UT[aidx] = T(tput[aidx] * servt[aidx]);
5693 hasUT[aidx] = true;
5694 // LINE scales the utilization of a queueing station into [0,1] whatever
5695 // its multiplicity, and reports a mean number of busy servers at an
5696 // infinite server: the processor share of an activity follows the same
5697 // convention
5698 const T hd = lqn.hostdem[aidx].disabled ? Tzero() : lqn.hostdem[aidx].mean;
5699 PN[aidx] = T(tput[aidx] * hd / num_traits<T>::from_double(ph_host_servers(hidx)));
5700 hasPN[aidx] = true;
5701 PN[hidx] = T(PN[hidx] + PN[aidx]);
5702 hasPN[hidx] = true;
5703 }
5704
5705 for (std::size_t e = 1; e <= lqn.nentries; ++e) {
5706 const std::size_t eidx = lqn.eshift + e;
5707 const std::size_t tidx = lqn.parent[eidx];
5708 if (ignore[tidx]) continue;
5709 out.TN[eidx] = tput[eidx];
5710 out.defined_T[eidx] = true;
5711 out.RN[eidx] = servt[eidx];
5712 out.defined_R[eidx] = true;
5713 UT[eidx] = T(tput[eidx] * servt[eidx]);
5714 hasUT[eidx] = true;
5715 for (std::size_t aidx : lqn.actsof[eidx]) {
5716 PN[eidx] = T(PN[eidx] + PN[aidx]);
5717 hasPN[eidx] = true;
5718 }
5719 // ResidT is reported per visit to the TASK, not per execution of the
5720 // activity: an activity of this entry runs EXECS times per invocation,
5721 // and the entry takes SHARE of the task's invocations. RespT stays per
5722 // execution.
5723 if (ph_has_wf[eidx]) {
5724 const std::unordered_map<std::size_t, T>& ex = ph_execs[eidx];
5725 for (std::size_t aidx : lqn.actsof[eidx]) {
5726 out.WN[aidx] = T(ph_share[eidx] * ex.at(aidx) * residt[aidx]);
5727 out.defined_W[aidx] = true;
5728 }
5729 }
5730 UT[tidx] = T(UT[tidx] + UT[eidx]);
5731 hasUT[tidx] = true;
5732 }
5733
5734 for (std::size_t t = 1; t <= lqn.ntasks; ++t) {
5735 const std::size_t tidx = lqn.tshift + t;
5736 if (ignore[tidx]) continue;
5737 out.TN[tidx] = tput[tidx];
5738 out.defined_T[tidx] = true;
5739 T w = Tzero();
5740 bool anyw = false;
5741 for (std::size_t aidx : lqn.actsof[tidx]) {
5742 PN[tidx] = T(PN[tidx] + PN[aidx]);
5743 hasPN[tidx] = true;
5744 if (out.defined_W[aidx]) { w = T(w + out.WN[aidx]); anyw = true; }
5745 }
5746 if (anyw) { out.WN[tidx] = w; out.defined_W[tidx] = true; }
5747 }
5748
5749 for (std::size_t hidx = 1; hidx <= lqn.nhosts; ++hidx)
5750 out.defined_T[hidx] = false; // kept undefined for consistency with LQNS
5751
5752 for (std::size_t idx = 1; idx <= N; ++idx) {
5753 out.QN[idx] = UT[idx];
5754 out.defined_Q[idx] = hasUT[idx];
5755 out.UN[idx] = PN[idx];
5756 out.defined_U[idx] = hasPN[idx];
5757 // Idle, not undefined -- the same rule aggregate() applies, and for the
5758 // same reason: an unreachable element reports zero for the measures its
5759 // kind HAS and leaves the ones it never has undefined, so that the
5760 // table's NaN mask survives a disconnected component.
5761 if (ignore[idx]) {
5762 out.UN[idx] = Tzero(); // every kind reports a utilization
5763 out.defined_U[idx] = true;
5764 out.defined_A[idx] = false; // nothing reports an arrival rate on an LQN
5765 const bool host = lqn.type[idx] == LqnElement::HOST;
5766 const bool task = lqn.type[idx] == LqnElement::TASK;
5767 const bool entry = lqn.type[idx] == LqnElement::ENTRY;
5768 out.QN[idx] = Tzero();
5769 out.defined_Q[idx] = !host;
5770 out.RN[idx] = Tzero();
5771 out.defined_R[idx] = !host && !task;
5772 out.WN[idx] = Tzero();
5773 out.defined_W[idx] = !host && !entry;
5774 out.TN[idx] = Tzero();
5775 out.defined_T[idx] = !host;
5776 }
5777 }
5778 return out;
5779 }
5780
5781 void update_metrics(int it) {
5782 if (is_ph_encoding()) {
5783 update_metrics_ph(it);
5784 return;
5785 }
5786 if (lnmethod == "moment3") {
5787 update_metrics_moment_based(it);
5788 return;
5789 }
5790 update_metrics_default(it);
5791 }
5792
5793 /**
5794 * Port of ln_layer_refcell: the 1-based class whose throughput at its
5795 * refstat normalises a residence of class K (0-based) in the layer of
5796 * element IDX, 0 when there is none. A chain merged by refpath holds more
5797 * than one caller task and only ONE reference class, so the per-caller rate
5798 * is read through the layer's normclass; everywhere else it is the chain's
5799 * reference class, which is what normclass would have named anyway.
5800 */
5801 std::size_t layer_refclass(std::size_t idx, const qn::Layer<T>& L, std::size_t k) const {
5802 const auto it = layer_normclass.find(idx);
5803 if (it != layer_normclass.end() && k < it->second.size() && it->second[k] > 0)
5804 return it->second[k];
5805 std::size_t c = 0;
5806 for (std::size_t cc = 0; cc < L.nchains; ++cc)
5807 if (L.chains[cc][k]) c = cc;
5808 return L.refclass[c];
5809 }
5810
5811 void update_metrics_default(int it) {
5812 const std::size_t N = lqn.nidx;
5813 servt.assign(N + 1, Tzero());
5814 residt.assign(N + 1, Tzero());
5815 const int iter_min = std::min(30, int(std::ceil(opt.iter_max / 4.0)));
5816 const bool averaging = averagingstart >= 0 && it >= iter_min;
5817 const int wnd = averaging ? int(it - averagingstart + 1) : 1;
5818
5819 for (const UpdRow& row : servt_map) {
5820 const std::size_t e = std::size_t(idxhash[row.idx]);
5821 const qn::Layer<T>& L = ensemble[e];
5822 const std::size_t k = row.cls - 1;
5823 const std::size_t refclass_c = layer_refclass(row.idx, L, k);
5824 const std::size_t refstat_k = L.classes[k].refstat;
5825
5826 T sv = Tzero(), rs = Tzero(), tp = Tzero();
5827 const T wT = T(Tone() / num_traits<T>::from_int(wnd));
5828 for (int w = 0; w < wnd; ++w) {
5829 const LayerResult<T>& r = results[results.size() - 1 - w][e];
5830 sv += r.RN(row.node - 1, k) * wT;
5831 const T TN_ref = (refclass_c > 0 && refstat_k > 0) ? r.TN(refstat_k - 1, refclass_c - 1) : Tzero();
5832 if (dbl(TN_ref) > GlobalConstants::FineTol)
5833 rs += r.QN(row.node - 1, k) / TN_ref * wT;
5834 else
5835 rs += r.WN(row.node - 1, k) * wT;
5836 tp += r.TN(row.node - 1, k) * wT;
5837 }
5838 servt[row.aidx] = sv;
5839 residt[row.aidx] = rs;
5840 tput[row.aidx] = tp;
5841
5842 // an activity think time is in series with the host demand
5843 const Distrib<T>& at = lqn.actthink[row.aidx];
5844 if (!at.disabled && dbl(at.mean) > GlobalConstants::FineTol) {
5845 servt[row.aidx] = T(servt[row.aidx] + at.mean);
5846 residt[row.aidx] = T(residt[row.aidx] + at.mean);
5847 }
5848
5849 // An activity of an async-only entry takes RN, the response per
5850 // visit, and not the visit-weighted residence: entry selection
5851 // routes the task to each of its entries with a share, and there is
5852 // no caller-side visit ratio here to divide that share back out
5853 // (the sync branch below does exactly that). updateMetricsDefault.m:63-77.
5854 if (async_only_activity(row.aidx)) residt[row.aidx] = servt[row.aidx];
5855
5856 if (relax_omega < 1.0 && it > 1) {
5857 const T om = num_traits<T>::from_double(relax_omega);
5858 const T om1 = num_traits<T>::from_double(1.0 - relax_omega);
5859 if (!std::isnan(servt_prev[row.aidx]))
5860 servt[row.aidx] = T(om * servt[row.aidx] + om1 * servt_prev_v[row.aidx]);
5861 if (!std::isnan(residt_prev[row.aidx]))
5862 residt[row.aidx] = T(om * residt[row.aidx] + om1 * residt_prev_v[row.aidx]);
5863 if (!std::isnan(tput_prev[row.aidx]))
5864 tput[row.aidx] = T(om * tput[row.aidx] + om1 * tput_prev_v[row.aidx]);
5865 }
5866 servt_prev[row.aidx] = dbl(servt[row.aidx]);
5867 residt_prev[row.aidx] = dbl(residt[row.aidx]);
5868 tput_prev[row.aidx] = dbl(tput[row.aidx]);
5869 servt_prev_v[row.aidx] = servt[row.aidx];
5870 residt_prev_v[row.aidx] = residt[row.aidx];
5871 tput_prev_v[row.aidx] = tput[row.aidx];
5872
5873 if (servt[row.aidx] > Tzero() && dbl(servt[row.aidx]) <= 1e10)
5874 servtproc[row.aidx] = Distrib<T>::exp_mean(servt[row.aidx]);
5875 tputproc[row.aidx] = Distrib<T>::exp_rate(tput[row.aidx]);
5876 }
5877
5878 // The phase split of servt, updateMetricsDefault.m:120-151. It is
5879 // recomputed from scratch each iteration because servt is; the
5880 // overtaking probability it feeds is computed later, once the entry
5881 // throughputs exist.
5882 if (has_phase2) {
5883 servt_ph1.assign(N + 1, Tzero());
5884 servt_ph2.assign(N + 1, Tzero());
5885 for (std::size_t a = 1; a <= lqn.nacts; ++a) {
5886 const std::size_t aidx = lqn.ashift + a;
5887 if (lqn.actphase[a] == 1)
5888 servt_ph1[aidx] = servt[aidx];
5889 else
5890 servt_ph2[aidx] = servt[aidx];
5891 }
5892 // the ENTRY split is taken once the entry service exists, below: summing the
5893 // activities' own servt left out the synchronous calls each phase makes
5894 }
5895
5896 // throughput of the activities that appear only as client-side classes
5897 for (const UpdRow& row : thinkt_map) {
5898 if (!tputproc[row.aidx].disabled) continue;
5899 const std::size_t e = std::size_t(idxhash[row.idx]);
5900 T tp = Tzero();
5901 const T wT = T(Tone() / num_traits<T>::from_int(wnd));
5902 for (int w = 0; w < wnd; ++w)
5903 tp += results[results.size() - 1 - w][e].TN(row.node - 1, row.cls - 1) * wT;
5904 tput[row.aidx] = tp;
5905 tputproc[row.aidx] = Distrib<T>::exp_rate(tp);
5906 }
5907
5908 // call service and residence times
5909 callservt.assign(lqn.ncalls + 1, Tzero());
5910 callresidt.assign(lqn.ncalls + 1, Tzero());
5911 for (const UpdRow& row : call_map) {
5912 if (row.node <= 1) continue; // a client-side call class contributes none
5913 const std::size_t e = std::size_t(idxhash[row.idx]);
5914 const LayerResult<T>& r = results.back()[e];
5915 callservt[row.aidx] =
5916 T(r.RN(row.node - 1, row.cls - 1) * lqn.callproc_mean[row.aidx]);
5917 // Normalise per chain-reference visit, as residt does. WN divides by the
5918 // class's own reference rate when the layer is open (an INF client task),
5919 // which is per-ENTRY visit, and the entry rescaling in
5920 // resolve_entry_service would then charge the call once per entry.
5921 {
5922 const qn::Layer<T>& L = ensemble[e];
5923 const std::size_t k = row.cls - 1;
5924 const std::size_t refclass_c = layer_refclass(row.idx, L, k);
5925 const std::size_t refstat_k = L.classes[k].refstat;
5926 const T TN_ref = (refclass_c > 0 && refstat_k > 0) ? r.TN(refstat_k - 1, refclass_c - 1) : Tzero();
5927 callresidt[row.aidx] = dbl(TN_ref) > GlobalConstants::FineTol
5928 ? T(r.QN(row.node - 1, k) / TN_ref)
5929 : r.WN(row.node - 1, k);
5930 }
5931 const T rw = region_wait(e, row.cls);
5932 if (rw > Tzero()) {
5933 callservt[row.aidx] = T(callservt[row.aidx] + rw);
5934 callresidt[row.aidx] = T(callresidt[row.aidx] + rw);
5935 }
5936 if (relax_omega < 1.0 && it > 1 && !std::isnan(callservt_prev[row.aidx])) {
5937 const T om = num_traits<T>::from_double(relax_omega);
5938 const T om1 = num_traits<T>::from_double(1.0 - relax_omega);
5939 callservt[row.aidx] =
5940 T(om * callservt[row.aidx] + om1 * callservt_prev_v[row.aidx]);
5941 }
5942 callservt_prev[row.aidx] = dbl(callservt[row.aidx]);
5943 callservt_prev_v[row.aidx] = callservt[row.aidx];
5944 callresidt_prev[row.aidx] = dbl(callresidt[row.aidx]);
5945 }
5946
5947 resolve_entry_service();
5948
5949 // The overtaking correction, updateMetricsDefault.m:313-351. It runs
5950 // HERE and not with the split above because it needs the entry
5951 // throughput resolve_entry_service has just produced.
5952 //
5953 // servt keeps both phases -- the server IS busy through phase 2, so the
5954 // utilization is unchanged -- while residt becomes the CALLER's view:
5955 // phase 1 in full, phase 2 only when the caller is actually overtaken.
5956 if (has_phase2) {
5957 // Entry phase split: the phase-2 share of the entry service, calls included,
5958 // scaled onto servt as resolve_entry_service scaled the total
5959 for (std::size_t e = 1; e <= lqn.nentries; ++e) {
5960 const std::size_t eidx = lqn.eshift + e;
5961 servt_ph2[eidx] = (eidx < entry_servt_last.size() && entry_servt_last[eidx] > Tzero())
5962 ? T(entry_servt_ph2[eidx] * servt[eidx] / entry_servt_last[eidx])
5963 : Tzero();
5964 servt_ph1[eidx] = T(servt[eidx] - servt_ph2[eidx]);
5965 }
5966 for (std::size_t e = 1; e <= lqn.nentries; ++e) {
5967 const std::size_t eidx = lqn.eshift + e;
5968 const std::size_t tidx = lqn.parent[eidx];
5969 if (!(dbl(servt_ph2[eidx]) > GlobalConstants::FineTol)) continue;
5970 if (lqn.isref[tidx] || !lqn.issynccaller.any_col(eidx)) {
5971 residt[eidx] = servt[eidx];
5972 continue;
5973 }
5974 T entry_tput = Tzero();
5975 if (dbl(tput[eidx]) > GlobalConstants::FineTol)
5976 entry_tput = tput[eidx];
5977 else if (dbl(tput[tidx]) > GlobalConstants::FineTol)
5978 entry_tput = tput[tidx];
5979 prOvertake[e] =
5980 dbl(entry_tput) > GlobalConstants::FineTol
5981 ? lqn_overtake_prob_markov(lqn, servt, callresidt, tput, eidx,
5982 servt_ph2[eidx])
5983 : Tzero();
5984 residt[eidx] = T(servt_ph1[eidx] + prOvertake[e] * servt_ph2[eidx]);
5985 }
5986 }
5987
5988 for (const UpdRow& row : call_map) {
5989 if (row.node <= 1) continue;
5990 const std::size_t eidx = lqn.callpair_dst[row.aidx];
5991 if (servt[eidx] > Tzero()) servtproc[eidx] = Distrib<T>::exp_mean(servt[eidx]);
5992 }
5993 for (const UpdRow& row : call_map) {
5994 if (row.node <= 1) continue;
5995 const std::size_t eidx = lqn.callpair_dst[row.aidx];
5996 if (it == 1) {
5997 callservt[row.aidx] = servt[eidx];
5998 callservtproc[row.aidx] = servtproc[eidx];
5999 } else if (callservt[row.aidx] > Tzero()) {
6000 callservtproc[row.aidx] = Distrib<T>::exp_mean(callservt[row.aidx]);
6001 }
6002 }
6003
6004 // What a synchronous CALLER waits for at a phase-2 target is residt,
6005 // not servt (updateMetricsDefault.m:384-400): the loop just above set it
6006 // from the target's full service, which would charge the caller for
6007 // phase 2 it never waits through.
6008 if (has_phase2) {
6009 for (std::size_t cidx = 1; cidx <= lqn.ncalls; ++cidx) {
6010 if (lqn.calltype[cidx] != CallType::SYNC) continue;
6011 const std::size_t target = lqn.callpair_dst[cidx];
6012 if (target <= lqn.eshift || target > lqn.eshift + lqn.nentries) continue;
6013 if (!(dbl(servt_ph2[target]) > GlobalConstants::FineTol)) continue;
6014 const T eff = residt[target]; // servt_ph1 + prOvertake * servt_ph2
6015 if (!(eff > Tzero())) continue;
6016 const T w = T(eff * lqn.callproc_mean[cidx]);
6017 callservt[cidx] = w;
6018 callresidt[cidx] = w;
6019 callservtproc[cidx] = Distrib<T>::exp_mean(w);
6020 }
6021 }
6022 }
6023
6024 // -----------------------------------------------------------------------
6025 // updateMetricsMomentBased, the `moment3` method
6026 // -----------------------------------------------------------------------
6027
6028 /**
6029 * The moments of an empirical CDF, MATLAB's `EmpiricalCDF.getMoments`.
6030 *
6031 * READ EXACTLY AS THE REFERENCE READS IT, and the reading is not the
6032 * obvious one: the weight of a bin is the CDF INCREMENT and its abscissa is
6033 * the MIDPOINT of the two t values, so this is a midpoint quadrature of
6034 * integral x^k dF and not a sum over grid points. Reading the grid as a pmf
6035 * at the right endpoint instead inflates every moment; that exact mistake
6036 * cost a wrong entry mean in the Python port (see _kb/06-solver-catalog.md).
6037 */
6038 static void cdf_moments(const fluid::FluidPassage& c, double& m1, double& m2, double& m3) {
6039 m1 = m2 = m3 = 0.0;
6040 for (std::size_t i = 0; i + 1 < c.t.size(); ++i) {
6041 const double x = 0.5 * (c.t[i + 1] + c.t[i]);
6042 const double w = c.cdf[i + 1] - c.cdf[i];
6043 m1 += x * w;
6044 m2 += x * x * w;
6045 m3 += x * x * x * w;
6046 }
6047 }
6048
6049 /** The per-layer response-time CDFs, computed once and cached. */
6050 const std::vector<std::vector<fluid::FluidPassage>>& layer_cdf(std::size_t e) {
6051 if (cdf_repo[e].empty()) {
6052 // THE LAYER HAS NO ROUTING MATRIX UNTIL IT IS ASKED FOR ONE, and the
6053 // fluid drift is built from `sn.rt` alone: without this the ODE has
6054 // no transitions, the state decays to zero, every passage reports the
6055 // degenerate curve at the origin and every entry ends up with an
6056 // empty convolution and a service time of zero. Same reason as in
6057 // solve_layer_fluid, and the failure is silent in both.
6058 ensemble[e].refresh_rt();
6059 try {
6060 cdf_repo[e] = detail::ln_fluid_cdf_respt(ensemble[e], opt.layer_fluid);
6061 } catch (const std::exception&) {
6062 // The reference falls back to the LAYER's own solver when the
6063 // fluid passage fails, and for the MVA-family layers that is
6064 // the base-class exponential law with the layer's mean -- so
6065 // the fallback here is that law over the last solved averages,
6066 // rather than an empty repo that silently collapses every
6067 // entry law to its bare host demand.
6068 cdf_repo[e] = layer_cdf_exp_fallback(e);
6069 }
6070 }
6071 return cdf_repo[e];
6072 }
6073
6074 /** The base-class exponential CDF over layer e's last solved mean response
6075 * times, the reference's fallback route when the fluid passage fails. */
6076 std::vector<std::vector<fluid::FluidPassage>> layer_cdf_exp_fallback(std::size_t e) {
6077 const qn::NetworkStruct<T>& L = ensemble[e];
6078 std::vector<std::vector<fluid::FluidPassage>> out(
6079 L.nstations, std::vector<fluid::FluidPassage>(L.nclasses));
6080 if (results.empty() || e >= results.back().size()) return out;
6081 const Matrix<T>& RN = results.back()[e].RN;
6082 const std::size_t npts = 100;
6083 for (std::size_t i = 0; i < L.nstations; ++i) {
6084 if (L.stations[i].nodetype == qn::NodeType::Source) continue;
6085 for (std::size_t r = 0; r < L.nclasses; ++r) {
6086 if (L.disabled[i][r]) continue;
6087 const double rn = (i < static_cast<std::size_t>(RN.rows()) &&
6088 r < static_cast<std::size_t>(RN.cols()))
6089 ? dbl(RN(i, r))
6090 : 0.0;
6091 if (!(std::isfinite(rn) && rn > 0.0)) continue;
6092 fluid::FluidPassage& cell = out[i][r];
6093 cell.t.reserve(npts);
6094 cell.cdf.reserve(npts);
6095 for (std::size_t j = 0; j < npts; ++j) {
6096 const double q =
6097 0.001 + (0.999 - 0.001) * static_cast<double>(j) / (npts - 1);
6098 cell.cdf.push_back(q);
6099 cell.t.push_back(-std::log(1.0 - q) * rn);
6100 }
6101 }
6102 }
6103 return out;
6104 }
6105
6106 /**
6107 * task_tput / entry_tput at the host layer of `eidx`.
6108 *
6109 * This is what renormalises a residence time from "one visit to the TASK",
6110 * which is how the layer reports it, to "one visit to the ENTRY", which is
6111 * what an entry metric means. It exists as a helper because applying it
6112 * TWICE is a real and silent failure mode: `moment3` once summed per-visit
6113 * response times and then applied this ratio on top, inflating every entry
6114 * service time by the entries-per-task ratio.
6115 *
6116 * `state` is 0 when the entry's layers are ignored (nothing is assigned at
6117 * all), 1 when it has no synchronous caller (no ratio exists) and 2 when
6118 * `ratio` is set.
6119 */
6120 int entry_visit_ratio(std::size_t eidx, T& ratio) const {
6121 const std::size_t tidx = lqn.parent[eidx];
6122 const std::size_t hidx = lqn.parent[tidx];
6123 if (ignore[tidx] || ignore[hidx] || idxhash[hidx] < 0) return 0;
6124 if (!lqn.issynccaller.any_col(eidx)) return 1;
6125 const std::size_t hl = std::size_t(idxhash[hidx]);
6126 const qn::Layer<T>& L = ensemble[hl];
6127 const LayerResult<T>& r = results.back()[hl];
6128 T task_tput = Tzero(), entry_tput = Tzero();
6129 for (const auto& kv : L.attr_tasks)
6130 if (kv.second == tidx) task_tput += r.TN(L.clientIdx - 1, kv.first - 1);
6131 for (const auto& kv : L.attr_entries)
6132 if (kv.second == eidx) entry_tput += r.TN(L.clientIdx - 1, kv.first - 1);
6133 const T floor = num_traits<T>::from_double(GlobalConstants::Zero);
6134 ratio = T(task_tput / (entry_tput > floor ? entry_tput : floor));
6135 return 2;
6136 }
6137
6138 /** (I - servtmatrix)^-1, the reference's `inv(eye - servtmatrix)`. */
6139 Matrix<T> entry_service_resolvent() const {
6140 const std::size_t dim = lqn.nidx + lqn.ncalls;
6141 Matrix<T> A(dim + 1, dim + 1, Tzero());
6142 for (std::size_t i = 0; i <= dim; ++i) {
6143 A(i, i) = Tone();
6144 for (std::size_t j = 0; j <= dim; ++j) A(i, j) = T(A(i, j) - servtmatrix(i, j));
6145 }
6146 return ::line::inverse(A);
6147 }
6148
6149 /**
6150 * Port of @@SolverLN/updateMetricsMomentBased.m, the `moment3` method.
6151 *
6152 * WHAT IT DOES DIFFERENTLY from the default update. The default feeds MEANS
6153 * between the layers. This fits an APH to the response-time CDF of every
6154 * activity and every call a layer reports, convolves those fits along the
6155 * entry's activity sequence, and reads the entry's law off the convolution:
6156 * a mean AND a distribution, which is what `get_cdf_respt` returns.
6157 *
6158 * TWO PASSES, not one. While the ensemble is still moving (`!hasconverged`)
6159 * the update is mean-based -- forming a CDF per layer per iteration would
6160 * cost a fluid integration per layer per iteration and would be fitting
6161 * noise anyway -- and the distribution is formed ONCE, on the converged
6162 * ensemble. The reference splits it exactly here.
6163 *
6164 * THE NORMALISATION TRAP. Both passes build the entry service from
6165 * RESIDENCE times, which are normalised to one visit to the TASK, and then
6166 * apply the task/entry throughput ratio, which converts that to one visit
6167 * to the ENTRY. Summing per-visit RESPONSE times and applying the ratio as
6168 * well applies the normalisation twice and inflates every multi-entry task
6169 * by its entries-per-task ratio. That was a live defect in all three
6170 * reference codebases until 2026-07-31.
6171 */
6172 void update_metrics_moment_based(int it) {
6173 const std::size_t N = lqn.nidx;
6174 servt.assign(N + 1, Tzero());
6175 residt.assign(N + 1, Tzero());
6176 callservt.assign(lqn.ncalls + 1, Tzero());
6177 callresidt.assign(lqn.ncalls + 1, Tzero());
6178
6179 // ---- what every activity's layer reports, common to both passes ----
6180 for (const UpdRow& row : servt_map) {
6181 const std::size_t e = std::size_t(idxhash[row.idx]);
6182 const qn::Layer<T>& L = ensemble[e];
6183 const LayerResult<T>& r = results.back()[e];
6184 const std::size_t k = row.cls - 1;
6185 std::size_t c = 0;
6186 for (std::size_t cc = 0; cc < L.nchains; ++cc)
6187 if (L.chains[cc][k]) c = cc;
6188 const std::size_t refclass_c = L.refclass[c];
6189 const std::size_t refstat_k = L.classes[k].refstat;
6190 const T TN_ref = (refclass_c > 0 && refstat_k > 0) ? r.TN(refstat_k - 1, refclass_c - 1) : Tzero();
6191
6192 tput[row.aidx] = r.TN(row.node - 1, k);
6193 residt[row.aidx] = dbl(TN_ref) > GlobalConstants::FineTol
6194 ? T(r.QN(row.node - 1, k) / TN_ref)
6195 : r.WN(row.node - 1, k);
6196 if (!hasconverged) {
6197 servt[row.aidx] = r.RN(row.node - 1, k);
6198 servtproc[row.aidx] = Distrib<T>::exp_mean(servt[row.aidx]);
6199 const Distrib<T>& at = lqn.actthink[row.aidx];
6200 if (!at.disabled && dbl(at.mean) > GlobalConstants::FineTol) {
6201 servt[row.aidx] = T(servt[row.aidx] + at.mean);
6202 residt[row.aidx] = T(residt[row.aidx] + at.mean);
6203 servtproc[row.aidx] = Distrib<T>::exp_mean(servt[row.aidx]);
6204 }
6205 // An activity of an async-only entry takes the per-visit
6206 // response, for the same reason as in the default update.
6207 if (async_only_activity(row.aidx)) residt[row.aidx] = servt[row.aidx];
6208 } else {
6209 servtcdf[row.aidx] = layer_cdf(e)[row.node - 1][k];
6210 }
6211 }
6212
6213 for (const UpdRow& row : call_map) {
6214 if (row.node <= 1) continue; // a client-side call class contributes none
6215 const std::size_t e = std::size_t(idxhash[row.idx]);
6216 const LayerResult<T>& r = results.back()[e];
6217 callresidt[row.aidx] = r.WN(row.node - 1, row.cls - 1);
6218 if (!hasconverged)
6219 callservt[row.aidx] =
6220 T(r.RN(row.node - 1, row.cls - 1) * lqn.callproc_mean[row.aidx]);
6221 else
6222 callservtcdf[row.aidx] = layer_cdf(e)[row.node - 1][row.cls - 1];
6223 }
6224
6225 if (!hasconverged)
6226 moment3_means_pass(it);
6227 else
6228 moment3_distribution_pass(it);
6229 }
6230
6231 /** The mean-based pass of moment3, run while the ensemble is still moving. */
6232 void moment3_means_pass(int it) {
6233 const std::size_t dim = lqn.nidx + lqn.ncalls;
6234 std::vector<T> x(dim + 1, Tzero());
6235 for (std::size_t i = 1; i <= lqn.nidx; ++i) x[i] = residt[i];
6236 for (std::size_t c = 1; c <= lqn.ncalls; ++c) x[lqn.nidx + c] = callresidt[c];
6237
6238 // entry_servt = (I - servtmatrix) \ [residt; callresidt]
6239 const Matrix<T> Rinv = entry_service_resolvent();
6240 std::vector<T> entry_servt(dim + 1, Tzero());
6241 for (std::size_t i = 1; i <= dim; ++i) {
6242 T s = Tzero();
6243 for (std::size_t j = 1; j <= dim; ++j)
6244 if (Rinv(i, j) != Tzero()) s += Rinv(i, j) * x[j];
6245 entry_servt[i] = s;
6246 }
6247 for (std::size_t i = 1; i <= lqn.eshift; ++i) entry_servt[i] = Tzero();
6248
6249 // NO forwarding propagation here. `lqn_fwd_rendezvous` has already
6250 // reconnected every forwarding chain reachable from a synchronous call to
6251 // the client that issued the rendezvous (Franks 1999, Sec. 3.3.1), so the
6252 // forwarded service is in the caller's chain before this runs; charging it
6253 // again inflated the caller by exactly the forwarded entry's mean. An
6254 // asynchronous call into a chain is left untouched there by design -- a
6255 // send-no-reply does not block -- so it must not accumulate the forwarded
6256 // service either. See BUGS.md BUG-91.
6257
6258 for (std::size_t e = 1; e <= lqn.nentries; ++e)
6259 servt[lqn.eshift + e] = entry_servt[lqn.eshift + e];
6260
6261 // entry_residt = servtmatrix * [residt; callresidt]
6262 std::vector<T> entry_residt(dim + 1, Tzero());
6263 for (std::size_t i = lqn.eshift + 1; i <= lqn.eshift + lqn.nentries; ++i) {
6264 T s = Tzero();
6265 for (std::size_t j = 1; j <= dim; ++j)
6266 if (servtmatrix(i, j) != Tzero()) s += servtmatrix(i, j) * x[j];
6267 entry_residt[i] = s;
6268 }
6269
6270 for (std::size_t e = 1; e <= lqn.nentries; ++e) {
6271 const std::size_t eidx = lqn.eshift + e;
6272 T ratio = Tone();
6273 const int state = entry_visit_ratio(eidx, ratio);
6274 if (state == 0) continue;
6275 if (state == 1) {
6276 residt[eidx] = entry_residt[eidx];
6277 continue;
6278 }
6279 servt[eidx] = T(entry_servt[eidx] * ratio);
6280 residt[eidx] = T(entry_residt[eidx] * ratio);
6281 }
6282
6283 for (const UpdRow& row : call_map) {
6284 if (row.node <= 1) continue;
6285 const std::size_t eidx = lqn.callpair_dst[row.aidx];
6286 if (servt[eidx] > Tzero()) servtproc[eidx] = Distrib<T>::exp_mean(servt[eidx]);
6287 }
6288 for (const UpdRow& row : call_map) {
6289 if (row.node <= 1) continue;
6290 const std::size_t eidx = lqn.callpair_dst[row.aidx];
6291 if (it == 1) {
6292 // A response time is per visit, so the number of calls is 1 here.
6293 callservt[row.aidx] = servt[eidx];
6294 callservtproc[row.aidx] = servtproc[eidx];
6295 } else if (callservt[row.aidx] > Tzero()) {
6296 callservtproc[row.aidx] = Distrib<T>::exp_mean(callservt[row.aidx]);
6297 }
6298 }
6299 }
6300
6301 /**
6302 * The distribution pass of moment3, run once on the converged ensemble.
6303 *
6304 * Every term reachable from an entry is fitted to an APH from the first
6305 * three moments of its response-time CDF, repeated as many times as the
6306 * entry-service matrix says it is visited, and the whole sequence is
6307 * convolved. A FRACTIONAL repetition count -- a call made 1.5 times on
6308 * average -- is realised as a branch between the fitted law and a point
6309 * mass at zero, which is the reference's `aph_simplify(..., pattern 3)`.
6310 */
6311 void moment3_distribution_pass(int it) {
6312 if constexpr (!num_traits<T>::has_transcendental) {
6313 throw UnsupportedError(
6314 "SolverLN: the 'moment3' method fits an APH to a response-time CDF, which needs "
6315 "square roots and a matrix exponential; rerun with --arith double or real");
6316 } else {
6317 const std::size_t dim = lqn.nidx + lqn.ncalls;
6318 const Matrix<T> Rinv = entry_service_resolvent();
6319
6320 // The point mass at zero the fractional branch mixes against.
6321 mam::AphPair<T> zero_law;
6322 zero_law.alpha.push_back(Tone());
6323 zero_law.S = Matrix<T>(1, 1, num_traits<T>::from_double(-GlobalConstants::Immediate));
6324
6325 for (std::size_t en = 1; en <= lqn.nentries; ++en) {
6326 const std::size_t eidx = lqn.eshift + en;
6327 std::vector<mam::AphPair<T>> seq;
6328 for (std::size_t fitidx = 1; fitidx <= dim; ++fitidx) {
6329 if (!(Rinv(eidx, fitidx) > Tzero())) continue;
6330 // Entries themselves are dropped: the entry's own service is
6331 // what is being assembled, and a host or task index carries
6332 // no response-time law of its own.
6333 if (fitidx <= lqn.eshift + lqn.nentries) continue;
6334 const bool is_call = fitidx > lqn.nidx;
6335 const fluid::FluidPassage& curve =
6336 is_call ? callservtcdf[fitidx - lqn.nidx] : servtcdf[fitidx];
6337 double m1 = 0.0, m2 = 0.0, m3 = 0.0;
6338 cdf_moments(curve, m1, m2, m3);
6339
6340 // An activity think time is in series with the host demand,
6341 // so its raw moments convolve with the measured ones.
6342 if (!is_call && !lqn.actthink[fitidx].disabled &&
6343 dbl(lqn.actthink[fitidx].mean) > GlobalConstants::FineTol) {
6344 const Distrib<T>& zt = lqn.actthink[fitidx];
6345 const double t1 = dbl(lang::dist_moment(zt, 1));
6346 const double t2 = dbl(lang::dist_moment(zt, 2));
6347 const double t3 = dbl(lang::dist_moment(zt, 3));
6348 m3 = m3 + 3.0 * m2 * t1 + 3.0 * m1 * t2 + t3;
6349 m2 = m2 + 2.0 * m1 * t1 + t2;
6350 m1 = m1 + t1;
6351 }
6352
6353 // CoarseTol and not FineTol: an Immediate activity has a
6354 // near-zero mean whose APH fit has rates of order 1e8, and
6355 // the matrix exponential of the convolution then does not
6356 // terminate. The reference skips those terms outright.
6357 if (!(m1 > GlobalConstants::CoarseTol)) continue;
6358
6359 const mam::AphFitResult<T> fit = mam::aph_fit(
6360 num_traits<T>::from_double(m1), num_traits<T>::from_double(m2),
6361 num_traits<T>::from_double(m3), 10u,
6362 num_traits<T>::from_double(GlobalConstants::FineTol));
6363 const mam::AphPair<T> law = aph_pair_of_map(fit.aph);
6364
6365 double reps = dbl(Rinv(eidx, fitidx));
6366 // servtmatrix carries 1.0 for a call; the mean NUMBER of
6367 // calls is what says how many times its law is convolved.
6368 if (is_call) reps *= dbl(lqn.callproc_mean[fitidx - lqn.nidx]);
6369 const long whole = static_cast<long>(std::floor(reps));
6370 const double frac = reps - static_cast<double>(whole);
6371 for (long q = 0; q < whole; ++q) seq.push_back(law);
6372 if (frac > 0.0)
6373 seq.push_back(mam::aph_simplify(law, zero_law,
6374 num_traits<T>::from_double(frac),
6375 num_traits<T>::from_double(1.0 - frac),
6377
6378 if (is_call) {
6379 const std::size_t cidx = fitidx - lqn.nidx;
6380 callservt[cidx] = num_traits<T>::from_double(m1);
6381 callservtproc[cidx] = Distrib<T>::exp_mean(callservt[cidx]);
6382 } else {
6383 servt[fitidx] = num_traits<T>::from_double(m1);
6384 servtproc[fitidx] = Distrib<T>::exp_mean(servt[fitidx]);
6385 }
6386 }
6387
6388 if (seq.empty()) {
6389 servt[eidx] = Tzero();
6390 continue;
6391 }
6392 const mam::AphPair<T> entry_law = mam::aph_convseq(seq);
6393 entryproc[en] = entry_law;
6394 servt[eidx] = Distrib<T>::ph_moment(entry_law.alpha, entry_law.S, 1);
6395 servtproc[eidx] = Distrib<T>::exp_mean(servt[eidx]);
6396 entrycdfrespt[en] = aph_eval_cdf(entry_law);
6397 }
6398
6399 // NO forwarding propagation here, for the reason given at the
6400 // entry_servt assembly above. See BUGS.md BUG-91.
6401
6402 // entry_residt = servtmatrix * [residt; callresidt], then the same
6403 // task/entry renormalisation the mean pass applies.
6404 std::vector<T> x(dim + 1, Tzero());
6405 for (std::size_t i = 1; i <= lqn.nidx; ++i) x[i] = residt[i];
6406 for (std::size_t c = 1; c <= lqn.ncalls; ++c) x[lqn.nidx + c] = callresidt[c];
6407 for (std::size_t en = 1; en <= lqn.nentries; ++en) {
6408 const std::size_t eidx = lqn.eshift + en;
6409 T s = Tzero();
6410 for (std::size_t j = 1; j <= dim; ++j)
6411 if (servtmatrix(eidx, j) != Tzero()) s += servtmatrix(eidx, j) * x[j];
6412 T ratio = Tone();
6413 const int state = entry_visit_ratio(eidx, ratio);
6414 if (state == 0) continue;
6415 residt[eidx] = state == 2 ? T(s * ratio) : s;
6416 }
6417
6418 if (it == 1)
6419 for (const UpdRow& row : call_map) {
6420 if (row.node <= 1) continue;
6421 const std::size_t eidx = lqn.callpair_dst[row.aidx];
6422 callservt[row.aidx] = servt[eidx];
6423 callservtproc[row.aidx] = Distrib<T>::exp_mean(servt[eidx]);
6424 }
6425
6426 // The entry servt is now the MEAN OF AN APH CONVOLUTION, not a sum
6427 // of residence times, which the interlock rescale in
6428 // `update_populations` has to know. See BUGS.md BUG-97.
6429 moment_pass_done = true;
6430 }
6431 }
6432
6433 /**
6434 * (alpha, S) of an APH handed back by aph_fit as a (D0, D1) pair.
6435 *
6436 * D1(i,j) = (-D0 e)_i alpha_j by construction, so alpha is any row of D1
6437 * divided by that row's exit rate; the first row with a positive exit rate
6438 * is taken. It is not recovered from the stationary phase distribution,
6439 * which is a different vector and would silently refit the law.
6440 */
6441 static mam::AphPair<T> aph_pair_of_map(const mam::Map<T>& m) {
6442 mam::AphPair<T> out;
6443 const std::size_t n = m.D0.rows();
6444 out.S = m.D0;
6445 out.alpha.assign(n, Tzero());
6446 for (std::size_t i = 0; i < n; ++i) {
6447 T rowsum = Tzero();
6448 for (std::size_t j = 0; j < n; ++j) rowsum += m.D1(i, j);
6449 if (!(dbl(rowsum) > GlobalConstants::Zero)) continue;
6450 for (std::size_t j = 0; j < n; ++j) out.alpha[j] = T(m.D1(i, j) / rowsum);
6451 return out;
6452 }
6453 // No phase can absorb: the law is degenerate, so it enters phase one and
6454 // stays there, which is what a zero alpha would NOT say.
6455 out.alpha[0] = Tone();
6456 return out;
6457 }
6458
6459 /**
6460 * F(t) = 1 - alpha exp(S t) e on the reference's grid, `APH.evalCDF` with
6461 * no argument: 500 points over [0, mean + 10 sigma].
6462 */
6463 static LnCdf aph_eval_cdf(const mam::AphPair<T>& law) {
6464 LnCdf out;
6465 const double m1 = dbl(Distrib<T>::ph_moment(law.alpha, law.S, 1));
6466 const double m2 = dbl(Distrib<T>::ph_moment(law.alpha, law.S, 2));
6467 const double var = m2 - m1 * m1;
6468 const double sigma = var > 0.0 ? std::sqrt(var) : 0.0;
6469 const double tmax = m1 + 10.0 * sigma;
6470 const std::size_t P = 500;
6471 out.t.resize(P);
6472 out.cdf.resize(P);
6473 for (std::size_t k = 0; k < P; ++k) {
6474 const double t = tmax * static_cast<double>(k) / static_cast<double>(P - 1);
6475 out.t[k] = t;
6476 Matrix<T> St(law.S.rows(), law.S.cols(), Tzero());
6477 for (std::size_t i = 0; i < law.S.rows(); ++i)
6478 for (std::size_t j = 0; j < law.S.cols(); ++j)
6479 St(i, j) = T(law.S(i, j) * num_traits<T>::from_double(t));
6480 const Matrix<T> E = ::line::expm(St);
6481 T surv = Tzero();
6482 for (std::size_t i = 0; i < E.rows(); ++i)
6483 for (std::size_t j = 0; j < E.cols(); ++j) surv += law.alpha[i] * E(i, j);
6484 out.cdf[k] = 1.0 - dbl(surv);
6485 }
6486 return out;
6487 }
6488
6489 /**
6490 * The AND-join completion times, and how much they undercut the serial sum.
6491 *
6492 * The branches of an AND fork run concurrently, so the time to clear the
6493 * join is the k-th smallest of the branch completion times, k being the
6494 * quorum. The entry-service reachability matrix cannot express that -- it
6495 * charges every activity of every branch to the entry, i.e. it serialises
6496 * them -- so the difference is recorded per join and applied as a
6497 * correction to any entry that reaches it.
6498 *
6499 * Branch times are taken as exponential, so the variance is the square of
6500 * the mean; that is the reference's assumption, not an approximation added
6501 * here.
6502 */
6503 std::vector<T> join_excess() const {
6504 std::vector<T> excess(lqn.nidx + 1, Tzero());
6505 std::vector<std::size_t> joined;
6506 for (std::size_t tail = 1; tail <= lqn.nidx; ++tail) {
6507 if (lqn.actpretype[tail] != PrecedenceType::PRE_AND) continue;
6508 for (std::size_t sx : lqn.graph.succ(tail))
6509 if (sx > lqn.ashift && sx <= lqn.ashift + lqn.nacts) joined.push_back(sx);
6510 }
6511 std::sort(joined.begin(), joined.end());
6512 joined.erase(std::unique(joined.begin(), joined.end()), joined.end());
6513 if (joined.empty()) return excess;
6514
6515 fj::LqnBranchView<T> view;
6516 view.graph = Matrix<T>(lqn.nidx, lqn.nidx, Tzero());
6517 for (std::size_t i = 1; i <= lqn.nidx; ++i)
6518 for (const auto& e : lqn.graph.row[i]) view.graph(i - 1, e.first - 1) = e.second;
6519 view.ashift = lqn.ashift;
6520 view.nacts = lqn.nacts;
6521 view.actposttype.assign(lqn.nidx, 0);
6522 for (std::size_t i = 1; i <= lqn.nidx; ++i)
6523 view.actposttype[i - 1] = static_cast<int>(lqn.actposttype[i]);
6524
6525 for (std::size_t aidx : joined) {
6526 const std::vector<std::vector<std::size_t>> branches =
6527 fj::fj_branch_members(view, aidx);
6528 if (branches.empty()) continue;
6529 std::vector<T> means;
6530 for (const auto& br : branches) {
6531 T s = Tzero();
6532 for (std::size_t a : br) {
6533 s += residt[a];
6534 // A branch activity with an Immediate host demand does all its work
6535 // in a rendezvous, so residt alone leaves this correction inert.
6536 if (a < lqn.callsof.size())
6537 for (std::size_t cidx : lqn.callsof[a])
6538 if (cidx < lqn.calltype.size() && lqn.calltype[cidx] == CallType::SYNC
6539 && cidx < callresidt.size())
6540 s += callresidt[cidx];
6541 }
6542 means.push_back(s);
6543 }
6544 if (means.size() == 1) continue; // a single branch cannot overlap
6545 std::size_t quorum = means.size();
6546 if (lqn.actquorum[aidx] >= 1 && lqn.actquorum[aidx] <= means.size())
6547 quorum = lqn.actquorum[aidx];
6548 std::vector<T> vars;
6549 for (const T& m2 : means) vars.push_back(T(m2 * m2));
6550 T serial = Tzero();
6551 for (const T& m2 : means) serial += m2;
6552 if constexpr (num_traits<T>::has_transcendental) {
6553 const fj::FJQuorumMomentsResult<T> q = fj::fj_quorum_moments(means, vars, quorum);
6554 excess[aidx] = T(q.m - serial);
6555 } else {
6556 throw UnsupportedError(
6557 "SolverLN: the AND-join completion time is a k-th order statistic fitted "
6558 "through a three-point distribution, which needs a square root; use the "
6559 "double or real backend for a model with an AND join");
6560 }
6561 }
6562 return excess;
6563 }
6564
6565 /** The block that turns activity residence times into entry service times. */
6566 void resolve_entry_service() {
6567 const std::size_t dim = lqn.nidx + lqn.ncalls;
6568 std::vector<T> x(dim + 1, Tzero());
6569 for (std::size_t i = 1; i <= lqn.nidx; ++i) x[i] = residt[i];
6570 for (std::size_t c = 1; c <= lqn.ncalls; ++c) x[lqn.nidx + c] = callresidt[c];
6571 std::vector<T> entry_servt(dim + 1, Tzero());
6572 for (std::size_t i = 1; i <= dim; ++i) {
6573 T s = Tzero();
6574 for (std::size_t j = 1; j <= dim; ++j)
6575 if (servtmatrix(i, j) != Tzero()) s += servtmatrix(i, j) * x[j];
6576 entry_servt[i] = s;
6577 }
6578
6579 // replace each reachable join's serialised branch sum by its concurrent
6580 // completion time
6581 const std::vector<T> excess = join_excess();
6582 for (std::size_t e = 1; e <= lqn.nentries; ++e) {
6583 const std::size_t eidx = lqn.eshift + e;
6584 for (std::size_t aidx = 1; aidx <= lqn.nidx; ++aidx) {
6585 if (excess[aidx] == Tzero()) continue;
6586 if (servtmatrix(eidx, aidx) > Tzero())
6587 entry_servt[eidx] = T(entry_servt[eidx] + excess[aidx]);
6588 }
6589 if (entry_servt[eidx] < Tzero()) entry_servt[eidx] = Tzero();
6590 }
6591
6592 // A SetupTask's cold start is charged HERE, to the entry, and with the
6593 // probability that the thread was actually found powered down. It is not
6594 // host demand, so it belongs to no activity's residence -- reporting it
6595 // there put RespT(A2) at 1.29479 against 0.333178 from LDES on lqn_setup,
6596 // the bare demand. The probability is the one 'srvn.ph' uses through
6597 // ph_setup_prob, so the two encodings charge the same thing
6598 // (updateMetricsDefault.m, lqn_setup_charge.m).
6599 for (std::size_t e = 1; e <= lqn.nentries; ++e) {
6600 const std::size_t eidx = lqn.eshift + e;
6601 const double c = setup_charge(lqn.parent[eidx]);
6602 if (c != 0.0)
6603 entry_servt[eidx] = T(entry_servt[eidx] + num_traits<T>::from_double(c));
6604 }
6605
6606 // The phase-2 part of each entry's service: phase-2 activities and the calls they
6607 // make, through the same reachability. The overtaking block splits servt with it.
6608 if (has_phase2) {
6609 std::vector<T> x2(dim + 1, Tzero());
6610 for (std::size_t a = 1; a <= lqn.nacts; ++a)
6611 if (lqn.actphase[a] > 1) x2[lqn.ashift + a] = x[lqn.ashift + a];
6612 for (std::size_t c = 1; c <= lqn.ncalls; ++c) {
6613 const std::size_t src = lqn.callpair_src[c];
6614 if (src > lqn.ashift && src <= lqn.ashift + lqn.nacts && lqn.actphase[src - lqn.ashift] > 1)
6615 x2[lqn.nidx + c] = x[lqn.nidx + c];
6616 }
6617 entry_servt_ph2.assign(dim + 1, Tzero());
6618 for (std::size_t i = lqn.eshift + 1; i <= dim; ++i) {
6619 T s2 = Tzero();
6620 for (std::size_t j = 1; j <= dim; ++j)
6621 if (servtmatrix(i, j) != Tzero()) s2 += servtmatrix(i, j) * x2[j];
6622 entry_servt_ph2[i] = s2;
6623 }
6624 entry_servt_last = entry_servt;
6625 }
6626
6627 // ResidT is normalised so the TASK has one visit; the entries need one
6628 // visit each, hence the throughput ratio below
6629 for (std::size_t e = 1; e <= lqn.nentries; ++e) {
6630 const std::size_t eidx = lqn.eshift + e;
6631 const std::size_t tidx = lqn.parent[eidx];
6632 const std::size_t hidx = lqn.parent[tidx];
6633 if (ignore[tidx] || ignore[hidx]) continue;
6634 if (idxhash[hidx] < 0) continue;
6635 const bool has_sync = lqn.issynccaller.any_col(eidx);
6636 if (!has_sync) {
6637 servt[eidx] = entry_servt[eidx];
6638 residt[eidx] = entry_servt[eidx];
6639 continue;
6640 }
6641 const std::size_t hl = std::size_t(idxhash[hidx]);
6642 const qn::Layer<T>& L = ensemble[hl];
6643 const LayerResult<T>& r = results.back()[hl];
6644 T task_tput = Tzero(), entry_tput = Tzero();
6645 for (const auto& kv : L.attr_tasks)
6646 if (kv.second == tidx) task_tput += r.TN(L.clientIdx - 1, kv.first - 1);
6647 for (const auto& kv : L.attr_entries)
6648 if (kv.second == eidx) entry_tput += r.TN(L.clientIdx - 1, kv.first - 1);
6649 if (dbl(entry_tput) > GlobalConstants::Zero) {
6650 servt[eidx] = T(entry_servt[eidx] * task_tput / entry_tput);
6651 residt[eidx] = servt[eidx];
6652 // the same ratio the other way round, the entry's invocations per
6653 // task invocation, which a refpath hop stage converts its call terms by
6654 if (dbl(task_tput) > GlobalConstants::Zero)
6655 entryvisits[eidx] = T(entry_tput / task_tput);
6656 } else {
6657 servt[eidx] = entry_servt[eidx];
6658 residt[eidx] = entry_servt[eidx];
6659 }
6660 }
6661 }
6662
6663 // -----------------------------------------------------------------------
6664 // updatePopulations (interlock correction)
6665 // -----------------------------------------------------------------------
6666 /**
6667 * True when the layer engine applies Eq. (4.7) inside its own MVA. A layer whose MVA path
6668 * has no interlock term would be moved to another algorithm by the matrix alone: exact
6669 * multiserver MVA would become AMVA, the product-form kernels would become the
6670 * load-dependent forward step. That swap is worth far more than the correction it carries,
6671 * and on a layer sitting near a bifurcation it turns the LN iteration into a limit cycle.
6672 * Such a layer keeps the residt scaling instead.
6673 */
6674 bool layer_takes_interlock(std::size_t e) const {
6675 return opt.layer_solver == "mva" && e < ensemble.size() &&
6676 mva::mva_carries_interlock(ensemble[e], opt.layer);
6677 }
6678
6679 /**
6680 * Class-level interlock matrix of one host layer. IL[r][s] is the share of the class-s
6681 * queue that a class-r arrival must not see at the host. It is kept CLASS-indexed, not
6682 * chain-indexed, so that a later chain refresh cannot leave it stale: the analyzer
6683 * aggregates it to chains against the struct it is about to solve. Two classes are
6684 * interlocked only if BOTH their tasks are, which is the 0/1 relation ir_mkj of Eq. (5);
6685 * the diagonal stays zero, since a request always sees its own class in full. The entry
6686 * is the Eq. (5) product Pr(IL_ms)*IR_ms*IR_mr, asymmetric in (r,s) because Pr(IL) is
6687 * taken from the QUEUED class s, so that the layer's ILw(r,s) = 1-IL(r,s) is the
6688 * lower-level adjustment rate r_lower. An empty result keeps the layer on the plain MVA
6689 * path.
6690 */
6691 std::vector<std::vector<double>> build_layer_interlock(std::size_t e,
6692 const std::vector<std::size_t>& hts,
6693 const std::vector<T>& prIL,
6694 const std::vector<T>& PrIL) const {
6695 const qn::Layer<T>& L = ensemble[e];
6696 const std::size_t R = L.nclasses;
6697 std::vector<double> cls_ir(R, 0.0); // IR
6698 std::vector<double> cls_pr(R, 0.0); // Pr(IL)
6699 auto stamp = [&](std::size_t classIdx, std::size_t tidx) {
6700 if (classIdx < 1 || classIdx > R) return;
6701 for (std::size_t i = 0; i < hts.size(); ++i)
6702 if (hts[i] == tidx) {
6703 cls_ir[classIdx - 1] = dbl(prIL[i]);
6704 cls_pr[classIdx - 1] = dbl(PrIL[i]);
6705 }
6706 };
6707 for (const auto& a : L.attr_tasks) stamp(a.first, a.second);
6708 for (const auto& a : L.attr_entries) stamp(a.first, lqn.parent[a.second]);
6709 for (const auto& a : L.attr_activities) stamp(a.first, lqn.parent[a.second]);
6710 for (const auto& a : L.attr_calls) stamp(a[0], lqn.parent[a[2]]);
6711
6712 std::vector<std::vector<double>> IL(R, std::vector<double>(R, 0.0));
6713 bool any = false;
6714 for (std::size_t r = 0; r < R; ++r) {
6715 if (!(cls_ir[r] > GlobalConstants::FineTol)) continue;
6716 for (std::size_t sIl = 0; sIl < R; ++sIl) {
6717 if (sIl == r || !(cls_ir[sIl] > GlobalConstants::FineTol)) continue;
6718 IL[r][sIl] = cls_pr[sIl] * cls_ir[sIl] * cls_ir[r];
6719 if (IL[r][sIl] > GlobalConstants::FineTol) any = true;
6720 }
6721 }
6722 return any ? IL : std::vector<std::vector<double>>();
6723 }
6724
6725 /**
6726 * Interlock probability for one (client, server) pair, as {IR, Pr(IL)} of Li and Franks,
6727 * "An improved interlocking correction for decomposition of layered queueing networks",
6728 * CCECE 2015, Eqs. (3) and (4). isProcessorHost selects the m' rule of lqns
6729 * Interlock::ilrate_pril_flow: at a PROCESSOR the common-source population is doubled
6730 * above 3 customers and squared at or below it, which is what turns m = 4 into the
6731 * pril = 1/8 its trace reports. The two factors are multiplied into the Eq. (5) rate by
6732 * build_layer_interlock, so neither carries the source count on its own -- that lives in
6733 * m'. This replaces the superseded (n_s-1)/n_s discount of Franks (1999), Eq. (4.7).
6734 */
6735 std::pair<T, T> interlock_prob(std::size_t client_tidx, std::size_t server_idx,
6736 bool isProcessorHost) const {
6737 const std::vector<std::size_t>& common = il_common_entries[server_idx];
6738 const double nsrc = il_num_sources[server_idx];
6739 if (nsrc == 0.0 || common.empty()) return std::make_pair(Tzero(), Tzero());
6740 const std::vector<std::size_t>& allsrc = il_src_all[server_idx];
6741 T sum_flow = Tzero();
6742 T sum_pril = Tzero();
6743 for (std::size_t ce : common) {
6744 const std::size_t srcTask = lqn.parent[ce];
6745 const std::size_t cen = ce - lqn.eshift;
6746 // population of this common source, in customer copies
6747 double m_src = lqn.mult[srcTask];
6748 if (!std::isfinite(m_src) || m_src < 1.0) m_src = 1.0;
6749 double m_eff = m_src;
6750 if (isProcessorHost) m_eff = (m_src > 3.0) ? (m_src + m_src) : (m_src * m_src);
6751 for (std::size_t dst : lqn.entriesof[client_tidx]) {
6752 const std::size_t dn = dst - lqn.eshift;
6753 if (dn < 1 || dn > lqn.nentries) continue;
6754 if (!(il_all(cen, dn) > Tzero())) continue;
6755 T ce_tput = tput[ce];
6756 if (!(dbl(ce_tput) > GlobalConstants::FineTol) && !lqn.actsof[ce].empty())
6757 ce_tput = tput[lqn.actsof[ce][0]];
6758 if (!(dbl(ce_tput) > GlobalConstants::FineTol)) ce_tput = tput[srcTask];
6759 if (!(dbl(ce_tput) > GlobalConstants::FineTol)) continue;
6760 // phase-2 entries are refused upstream, so every source is all-phase
6761 if (std::find(allsrc.begin(), allsrc.end(), srcTask) != allsrc.end()) {
6762 const T contrib = T(ce_tput * il_all(cen, dn));
6763 sum_flow += contrib;
6764 sum_pril += T(contrib / num_traits<T>::from_double(m_eff));
6765 }
6766 }
6767 }
6768 T client_tput = tput[client_tidx];
6769 if (!(dbl(client_tput) > GlobalConstants::FineTol)) {
6770 for (std::size_t e : lqn.entriesof[client_tidx]) {
6771 T et = tput[e];
6772 if (!(dbl(et) > GlobalConstants::FineTol) && !lqn.actsof[e].empty())
6773 et = tput[lqn.actsof[e][0]];
6774 client_tput = T(client_tput + et);
6775 }
6776 }
6777 if (!(dbl(client_tput) > GlobalConstants::FineTol))
6778 return std::make_pair(Tzero(), Tzero());
6779 const T flow = sum_flow < client_tput ? sum_flow : client_tput;
6780 T IR = T(flow / client_tput);
6781 if (IR > Tone()) IR = Tone();
6782 if (dbl(IR) < 0.0) IR = Tzero();
6783 T pr = Tzero();
6784 if (dbl(sum_flow) > GlobalConstants::FineTol) pr = T(sum_pril / sum_flow);
6785 if (pr > Tone()) pr = Tone();
6786 if (dbl(pr) < 0.0) pr = Tzero();
6787 return std::make_pair(IR, pr);
6788 }
6789
6790 void update_populations(int) {
6791 const std::vector<T> residt_orig = residt;
6792 const std::vector<T> callresidt_orig = callresidt;
6793 bool adjusted = false;
6794
6795 for (std::size_t cidx = 1; cidx <= lqn.ncalls; ++cidx) {
6796 if (lqn.calltype[cidx] != CallType::SYNC) continue;
6797 const std::size_t dst = lqn.callpair_dst[cidx];
6798 const std::size_t server_tidx = lqn.parent[dst];
6799 std::size_t server = 0;
6800 if (server_tidx <= NT() && !il_common_entries[server_tidx].empty()) {
6801 server = server_tidx;
6802 } else if (server_tidx > lqn.tshift) {
6803 const std::size_t h = lqn.parent[server_tidx];
6804 if (h >= 1 && h <= NT() && !il_common_entries[h].empty()) server = h;
6805 }
6806 if (server == 0) continue;
6807 const std::size_t client_tidx = lqn.parent[lqn.callpair_src[cidx]];
6808 // This path serves a TASK, not a processor, so the m' rule of Li/lqns leaves the
6809 // source population alone; the product IR*Pr(IL) reproduces the scalar this
6810 // branch used before.
6811 const std::pair<T, T> ilp = interlock_prob(client_tidx, server, false);
6812 const T prIL = T(ilp.first * ilp.second);
6813 if (!(dbl(prIL) > GlobalConstants::FineTol)) continue;
6814 const T S = servt[dst];
6815 const T cm = lqn.callproc_mean[cidx];
6816 if (!(cm > Tzero()) || !(callservt[cidx] > Tzero())) continue;
6817 const T RN = T(callservt[cidx] / cm);
6818 const T W = RN > S ? T(RN - S) : Tzero();
6819 if (!(dbl(W) > GlobalConstants::FineTol)) continue;
6820 const T RN_adj = T(S + (Tone() - prIL) * W);
6821 const T scale = T(RN_adj / RN);
6822 callservt[cidx] = T(callservt[cidx] * scale);
6823 callresidt[cidx] = T(callresidt[cidx] * scale);
6824 if (callservt[cidx] > Tzero())
6825 callservtproc[cidx] = Distrib<T>::exp_mean(callservt[cidx]);
6826 adjusted = true;
6827 }
6828
6829 // Every layer starts the pass without a matrix, so a host that stops being
6830 // interlocked does not keep the previous iteration's correction alive.
6831 layer_interlock.assign(ensemble.size(), {});
6832 for (std::size_t h = 1; h <= lqn.nhosts; ++h) {
6833 if (il_common_entries[h].empty()) continue;
6834 const std::vector<std::size_t>& hts = lqn.tasksof[h];
6835 std::vector<T> prIL(hts.size(), Tzero()), PrIL(hts.size(), Tzero()),
6836 tutil(hts.size(), Tzero());
6837 for (std::size_t i = 0; i < hts.size(); ++i) {
6838 // The host of a task layer is a PROCESSOR, which is what selects the m' rule.
6839 const std::pair<T, T> ilp = interlock_prob(hts[i], h, true);
6840 prIL[i] = ilp.first;
6841 PrIL[i] = ilp.second;
6842 for (std::size_t e : lqn.entriesof[hts[i]])
6843 for (std::size_t a : lqn.actsof[e]) tutil[i] += tput[a] * lqn.hostdem[a].mean;
6844 }
6845 T Utot = Tzero(), Uil = Tzero();
6846 for (std::size_t i = 0; i < hts.size(); ++i) {
6847 Utot += tutil[i];
6848 if (dbl(prIL[i]) > GlobalConstants::FineTol) Uil += tutil[i];
6849 }
6850 if (!(dbl(Utot) > GlobalConstants::FineTol) || !(dbl(Uil) > GlobalConstants::FineTol))
6851 continue;
6852 const T frac = T(Uil / Utot);
6853
6854 // When the layer engine carries Eq. (4.7) inside its own MVA, the interlock goes
6855 // to the layer as a class-level matrix and the residence times are left
6856 // untouched. Scaling them here as well would remove the same waiting twice, and
6857 // would still leave the layer's own THROUGHPUT uncorrected, which is what breaks
6858 // flow balance across a call: the reported task rate then comes from a cycle time
6859 // the correction has already shortened elsewhere.
6860 if (idxhash[h] >= 0 && layer_takes_interlock(std::size_t(idxhash[h]))) {
6861 const std::size_t e = std::size_t(idxhash[h]);
6862 layer_interlock[e] = build_layer_interlock(e, hts, prIL, PrIL);
6863 continue;
6864 }
6865
6866 for (std::size_t i = 0; i < hts.size(); ++i) {
6867 if (!(dbl(prIL[i]) > GlobalConstants::FineTol)) continue;
6868 // Weight by the share of host utilization that is interlocked. The rate
6869 // is the SAME Eq. (5) product IR*Pr(IL) that pass 1 applies to a call and
6870 // that build_layer_interlock puts in the layer matrix -- IR alone is a
6871 // flow SHARE, ~1 whenever a layer has a single common source, and using
6872 // it here removed the whole processor queueing rather than the
6873 // interlocked part of it, which broke flow balance across a call.
6874 const T eff = T(prIL[i] * PrIL[i] * frac);
6875 for (std::size_t e : lqn.entriesof[hts[i]])
6876 for (std::size_t a : lqn.actsof[e]) {
6877 const T D = lqn.hostdem[a].mean;
6878 if (D > Tzero() && dbl(residt[a] - D) > GlobalConstants::FineTol) {
6879 residt[a] = T(D + (Tone() - eff) * (residt[a] - D));
6880 adjusted = true;
6881 }
6882 }
6883 }
6884 }
6885
6886 if (!adjusted) return;
6887 const std::size_t dim = lqn.nidx + lqn.ncalls;
6888 auto esum = [&](const std::vector<T>& rs, const std::vector<T>& cr, std::size_t i) {
6889 T s = Tzero();
6890 for (std::size_t j = 1; j <= dim; ++j) {
6891 if (servtmatrix(i, j) == Tzero()) continue;
6892 s += servtmatrix(i, j) * (j <= lqn.nidx ? rs[j] : cr[j - lqn.nidx]);
6893 }
6894 return s;
6895 };
6896 for (std::size_t e = 1; e <= lqn.nentries; ++e) {
6897 const std::size_t eidx = lqn.eshift + e;
6898 const T oldv = esum(residt_orig, callresidt_orig, eidx);
6899 if (!(dbl(oldv) > GlobalConstants::FineTol)) continue;
6900 const T ratio = T(esum(residt, callresidt, eidx) / oldv);
6901 // The entry servt is rescaled only when it was itself assembled from
6902 // these residence times, which is the mean-based path. After the
6903 // moment3 distribution pass it is the mean of an APH convolution of
6904 // the activities' own response laws, and a ratio of residence-time
6905 // sums is not a correction to it: applying it multiplied the entry
6906 // law by the entry's visit ratio and reported a service time BELOW
6907 // that of the single activity the entry contains. The residence
6908 // times keep their correction either way. See BUGS.md BUG-97.
6909 if (!moment_pass_done) {
6910 servt[eidx] = T(servt[eidx] * ratio);
6911 if (servt[eidx] > Tzero()) servtproc[eidx] = Distrib<T>::exp_mean(servt[eidx]);
6912 }
6913 residt[eidx] = T(residt[eidx] * ratio);
6914 }
6915 }
6916
6917 // -----------------------------------------------------------------------
6918 // Declared think time of a task as it enters the thread cycle: the value for
6919 // a REFERENCE task, zero for any other.
6920 //
6921 // A think time is an attribute of the closed customer population a reference
6922 // task stands for, and it is what separates one request of that population
6923 // from the next. On a served task it has no such meaning, and charging it
6924 // per request throttles the task: lqn_basic's T3 has 25 threads and a
6925 // declared think time of 4, and reading it as a per-request delay caps it at
6926 // 25/(4+0.02) = 6.219 completions per second. Three independent oracles put
6927 // the rate at five calls per caller request instead -- lqsim 66.5, LDES
6928 // 66.955, lqns 75.6. See _kb/06-solver-catalog.md (LN section).
6929 T ref_think_time(std::size_t tidx) const {
6930 if (tidx >= lqn.isref.size() || !lqn.isref[tidx]) return Tzero();
6931 if (lqn.think[tidx].disabled) return Tzero();
6932 return lqn.think[tidx].mean;
6933 }
6934
6935 // updateThinkTimes
6936 // -----------------------------------------------------------------------
6937 void update_think_times(int it) {
6938 // Under "srvn.ph" a caller reaches the server once per invocation, so the
6939 // station rate is not the task's invocation rate
6940 if (is_ph_encoding()) {
6941 update_think_times_ph(it);
6942 return;
6943 }
6944 thinktproc.assign(lqn.nidx + 1, Distrib<T>::disabled_dist());
6945 for (std::size_t t = 1; t <= lqn.ntasks; ++t) {
6946 const std::size_t tidx = lqn.tshift + t;
6947 // only a REFERENCE task's think time is a per-request delay
6948 const T ztask = ref_think_time(tidx);
6949 if (idxhash[tidx] < 0) {
6950 // A task reached only by an entry arrival still has a cycle: its threads
6951 // are driven by the stream. build_layer drops the open class for it so
6952 // the chain can be closed on the known rate here.
6953 const double arvrate = open_arrival_rate_of(tidx);
6954 if (arvrate > GlobalConstants::FineTol) {
6955 double nja = 0.0;
6956 for (std::size_t c = 1; c <= NT(); ++c) nja = std::max(nja, njobs(tidx, c));
6957 if (!(nja > 0.0)) nja = lqn.maxmult[tidx];
6958 T hres = Tzero();
6959 for (std::size_t eidx : lqn.entriesof[tidx])
6960 for (std::size_t aidx : lqn.actsof[eidx])
6961 if (!std::isnan(dbl(residt[aidx]))) hres += residt[aidx];
6962 const T floora = num_traits<T>::from_double(GlobalConstants::Zero);
6963 T za = T(num_traits<T>::from_double(nja / arvrate) - hres - ztask);
6964 if (za < floora) za = floora;
6965 if (relax_omega < 1.0 && it > 1 && !std::isnan(thinkt_prev[tidx])) {
6966 const T om = num_traits<T>::from_double(relax_omega);
6967 const T om1 = num_traits<T>::from_double(1.0 - relax_omega);
6968 za = T(om * za + om1 * thinkt_prev_v[tidx]);
6969 }
6970 tput[tidx] = num_traits<T>::from_double(arvrate);
6971 thinkt[tidx] = za;
6972 thinkt_prev[tidx] = dbl(za);
6973 thinkt_prev_v[tidx] = za;
6974 thinktproc[tidx] = Distrib<T>::exp_mean(T(za + ztask));
6975 continue;
6976 }
6977 thinkt[tidx] = num_traits<T>::from_double(GlobalConstants::FineTol);
6978 thinktproc[tidx] = Distrib<T>::immediate();
6979 continue;
6980 }
6981 const std::size_t e = std::size_t(idxhash[tidx]);
6982 const qn::Layer<T>& L = ensemble[e];
6983 const LayerResult<T>& r = results.back()[e];
6984 double nj = 0.0;
6985 for (std::size_t c = 1; c <= NT(); ++c) nj = std::max(nj, njobs(tidx, c));
6986 // the task's OWN station, which under `flat.cs` is one of many
6987 const std::size_t ts = station_idx_of(L, tidx);
6988 T tp = Tzero(), ut = Tzero();
6989 for (std::size_t k = 0; k < L.nclasses; ++k) {
6990 tp += r.TN(ts - 1, k);
6991 ut += r.UN(ts - 1, k);
6992 }
6993 tput[tidx] = T(num_traits<T>::from_double(lqn.repl[tidx]) * tp);
6994 util[tidx] = ut;
6995 T raw;
6996 if (lqn.sched[tidx] == SchedStrategy::INF) {
6997 // an infinite server reports utilization as a mean job count
6998 raw = tput[tidx] == Tzero()
6999 ? Tzero()
7000 : T((num_traits<T>::from_double(nj) - util[tidx]) / tput[tidx] - ztask);
7001 } else {
7002 const T om = util[tidx] > Tone() ? T(util[tidx] - Tone()) : T(Tone() - util[tidx]);
7003 raw = tput[tidx] == Tzero()
7004 ? Tzero()
7005 : T(num_traits<T>::from_double(nj) * om / tput[tidx] - ztask);
7006 }
7007 const T floorv = num_traits<T>::from_double(GlobalConstants::Zero);
7008 thinkt[tidx] = raw < floorv ? floorv : raw;
7009
7010 // The cold start goes the OTHER way from the phase-2 tail. A caller
7011 // class cycles as delay plus station service, and the station serves
7012 // only the host demand: the charge is on the ENTRY, not on any
7013 // activity's demand, so the station never sees it and the delay has
7014 // to carry it. Without this the callee layer cycles faster than its
7015 // callers drive it -- 0.529412 against 0.5 on lqn_setup, with the
7016 // caller conserved and the callee not. Zero for a task with no setup
7017 // (updateThinkTimes.m).
7018 {
7019 const double c = setup_charge(tidx);
7020 if (c != 0.0) {
7021 const T withc = T(thinkt[tidx] + num_traits<T>::from_double(c));
7022 thinkt[tidx] = withc < floorv ? floorv : withc;
7023 }
7024 }
7025
7026 if (relax_omega < 1.0 && it > 1 && !std::isnan(thinkt_prev[tidx])) {
7027 const double rawd = dbl(thinkt[tidx]);
7028 if (thinkt_prev[tidx] > 10.0 * rawd && rawd > GlobalConstants::FineTol) {
7029 thinkt_prev[tidx] = rawd;
7030 thinkt_prev_v[tidx] = thinkt[tidx];
7031 }
7032 const T om = num_traits<T>::from_double(relax_omega);
7033 const T om1 = num_traits<T>::from_double(1.0 - relax_omega);
7034 thinkt[tidx] = T(om * thinkt[tidx] + om1 * thinkt_prev_v[tidx]);
7035 }
7036 thinkt_prev[tidx] = dbl(thinkt[tidx]);
7037 thinkt_prev_v[tidx] = thinkt[tidx];
7038 thinktproc[tidx] = Distrib<T>::exp_mean(T(thinkt[tidx] + ztask));
7039 }
7040 }
7041
7042 // -----------------------------------------------------------------------
7043 // updateLayers and updateRoutingProbabilities
7044 // -----------------------------------------------------------------------
7045 void update_layers(int it) {
7046 // Under "srvn.ph" the layer classes are one per caller task and their laws
7047 // are composed, not read off the update maps
7048 if (is_ph_encoding()) {
7049 update_layers_ph(it);
7050 return;
7051 }
7052 const bool elevator = (it % 2) == 1;
7053 const std::size_t nt = thinkt_map.size();
7054 for (std::size_t r = 0; r < nt; ++r) {
7055 const UpdRow& row = thinkt_map[elevator ? nt - 1 - r : r];
7056 qn::Layer<T>& L = ensemble[std::size_t(idxhash[row.idx])];
7057 if (row.aidx <= NT() && opt.interlocking &&
7058 L.classes[row.cls - 1].type == JobClassType::CLOSED)
7059 L.classes[row.cls - 1].population = njobs(row.aidx, row.idx);
7060 if (row.node == L.clientIdx) {
7061 if (lqn.type[row.aidx] == LqnElement::TASK) {
7062 if (lqn.sched[row.aidx] != SchedStrategy::REF) {
7063 if (!thinktproc[row.aidx].disabled)
7064 L.set_service(row.node, row.cls, thinktproc[row.aidx]);
7065 } else {
7066 L.set_service(row.node, row.cls, servtproc[row.aidx]);
7067 }
7068 } else {
7069 L.set_service(row.node, row.cls, servtproc[row.aidx]);
7070 }
7071 } else {
7072 L.set_service(row.node, row.cls, servtproc[row.aidx]);
7073 }
7074 }
7075 const std::size_t nc = call_map.size();
7076 for (std::size_t r = 0; r < nc; ++r) {
7077 const UpdRow& row = call_map[elevator ? nc - 1 - r : r];
7078 qn::Layer<T>& L = ensemble[std::size_t(idxhash[row.idx])];
7079 if (row.node == L.clientIdx)
7080 L.set_service(row.node, row.cls, callservtproc[row.aidx]);
7081 else
7082 L.set_service(row.node, row.cls, servtproc[lqn.callpair_dst[row.aidx]]);
7083 }
7084 // Async arrival rate: the caller activity fires at its own throughput,
7085 // and every firing releases callmean jobs into this layer's Source.
7086 // updateLayers.m:63-73 replays the same map; the geometric self-loop in
7087 // build_layer already accounts for callmean, so the rate is the bare
7088 // activity throughput, not scaled by it.
7089 //
7090 // The reference-path stage means go first, after the think times:
7091 // residt, callresidt, util and tput are all fresh by now.
7092 if (interlockMethod == "refpath") update_ref_path_stages(it);
7093 for (const UpdRow& row : arv_call_map) {
7094 qn::Layer<T>& L = ensemble[std::size_t(idxhash[row.idx])];
7095 const Distrib<T>& d = tputproc[lqn.callpair_src[row.aidx]];
7096 if (!d.disabled) L.set_service(row.node, row.cls, d);
7097 }
7098 }
7099
7100 /**
7101 * Port of updateRefPathStages: recompute every reference-path stage mean.
7102 *
7103 * A HOP stage stands for an entry whose task has no activity graph in the
7104 * layer: residt(e), minus callresidt(c)/entryvisits(e) for each descent the
7105 * chain re-expands below it, plus the wait W for one of its task's threads.
7106 * A GATE into a MEMBER charges W(callee task) alone, since everything else
7107 * the member owns is explicit in the layer. W is taken in the POPULATION
7108 * domain, never as the difference of two nearly equal response times.
7109 */
7110 void update_ref_path_stages(int it) {
7111 if (refpath_map.empty()) return;
7112 const std::size_t nkeys = refpath_keys.size();
7113 std::vector<T> stage(nkeys, Tzero());
7114 for (std::size_t r = 0; r < refpath_map.size(); ++r) {
7115 const RefPathRow& row = refpath_map[r];
7116 T term = Tzero();
7117 switch (row.kind) {
7118 case 1: // the entry's own residence, per entry invocation
7119 term = residt[row.elem];
7120 break;
7121 case 2: { // a re-expanded descent, converted to the entry scale
7122 T v = Tone();
7123 if (row.fromentry >= 1 && dbl(entryvisits[row.fromentry]) > GlobalConstants::FineTol)
7124 v = entryvisits[row.fromentry];
7125 term = T(callresidt[row.elem] / v);
7126 break;
7127 }
7128 case 3: // the thread-acquisition wait of the task
7129 term = thread_wait(row.elem);
7130 break;
7131 }
7132 if (!std::isfinite(dbl(term))) term = Tzero();
7133 stage[refpath_group[r]] = T(stage[refpath_group[r]] + term * num_traits<T>::from_double(row.sign));
7134 }
7135 const double omega = relax_omega;
7136 for (std::size_t k = 0; k < nkeys; ++k) {
7137 if (stage[k] < Tzero()) stage[k] = Tzero();
7138 // the think-time schedule, including the escape from a divergent iterate
7139 if (omega < 1.0 && it > 1 && !std::isnan(refpath_stage_prev[k])) {
7140 T prevS = refpath_stage_prev_v[k];
7141 if (dbl(prevS) > 10.0 * dbl(stage[k]) && dbl(stage[k]) > GlobalConstants::FineTol)
7142 prevS = stage[k];
7143 stage[k] = T(num_traits<T>::from_double(omega) * stage[k] +
7144 num_traits<T>::from_double(1.0 - omega) * prevS);
7145 }
7146 refpath_stage_prev[k] = dbl(stage[k]);
7147 refpath_stage_prev_v[k] = stage[k];
7148 qn::Layer<T>& L = ensemble[std::size_t(idxhash[refpath_keys[k][0]])];
7149 // exp_mean(0) is an infinite rate, which the visit solve reads as a
7150 // disabled pair and drops from the chain
7151 if (dbl(stage[k]) > GlobalConstants::FineTol)
7152 L.set_service(refpath_keys[k][1], refpath_keys[k][2], Distrib<T>::exp_mean(stage[k]));
7153 else
7154 L.set_service(refpath_keys[k][1], refpath_keys[k][2], Distrib<T>::immediate());
7155 }
7156 }
7157
7158 /**
7159 * Mean time a request queues for a thread of task TIDX, read from TIDX's own
7160 * layer in the population domain, max(0, Q - m*U)/X. An infinite or
7161 * unbounded pool never queues.
7162 */
7163 T thread_wait(std::size_t tidx) const {
7164 if (tidx < 1 || tidx > lqn.nidx) return Tzero();
7165 if (lqn.sched[tidx] == SchedStrategy::INF || !std::isfinite(lqn.maxmult[tidx])) return Tzero();
7166 if (idxhash[tidx] < 0 || results.empty()) return Tzero();
7167 const std::size_t e = std::size_t(idxhash[tidx]);
7168 const qn::Layer<T>& L = ensemble[e];
7169 const LayerResult<T>& r = results.back()[e];
7170 const std::size_t ks = station_idx_of(L, tidx);
7171 T Qtot = Tzero();
7172 for (std::size_t c = 0; c < L.nclasses; ++c) Qtot += r.QN(ks - 1, c);
7173 const T Btot = T(num_traits<T>::from_double(lqn.maxmult[tidx]) * util[tidx]);
7174 const T Xtot = T(tput[tidx] / num_traits<T>::from_double(std::max(1.0, lqn.repl[tidx])));
7175 if (!(dbl(Xtot) > GlobalConstants::FineTol)) return Tzero();
7176 T d = T(Qtot - Btot);
7177 const T z = num_traits<T>::from_double(GlobalConstants::Zero);
7178 if (d < z) d = z;
7179 return T(d / Xtot);
7180 }
7181
7182 /**
7183 * Time a job of `cls`'s CHAIN spends blocked at this layer's region.
7184 *
7185 * A job held by an admission constraint is at no station at all, so its
7186 * wait is structurally absent from the RN and WN the layer reports back and
7187 * the caller and the callee end up disagreeing on throughput -- the fixed
7188 * point still converges, it just converges to the wrong place.
7189 *
7190 * Recovered by Little over the CHAIN, never per class: a job switches class
7191 * along the activity graph, so a call class carries population 0 and a
7192 * per-class deficit comes out negative and silently does nothing.
7193 * Zero for a layer with no region, which is every layer in a plain model.
7194 */
7195 T region_wait(std::size_t e, std::size_t cls) const {
7196 const qn::Layer<T>& L = ensemble[e];
7197 if (L.regions.empty() || results.empty()) return Tzero();
7198 const LayerResult<T>& r = results.back()[e];
7199 const std::size_t k = cls - 1;
7200 std::size_t c = L.nchains;
7201 for (std::size_t cc = 0; cc < L.nchains; ++cc)
7202 if (L.chains[cc][k]) c = cc;
7203 if (c == L.nchains) return Tzero();
7204
7205 // the station the region is stated over, which under `flat.cs` is one
7206 // among many and under `srvn` is the layer's own server
7207 std::size_t rs = L.serverIdx;
7208 for (std::size_t i = 0; i < L.regions[0].members.size(); ++i)
7209 if (L.regions[0].members[i]) { rs = i + 1; break; }
7210
7211 double pop = 0.0;
7212 T inside = Tzero(), tput_srv = Tzero();
7213 for (std::size_t j = 0; j < L.nclasses; ++j) {
7214 if (!L.chains[c][j]) continue;
7215 const double p = L.classes[j].population;
7216 if (!std::isfinite(p)) return Tzero(); // an open chain has no population to close on
7217 pop += p;
7218 for (std::size_t i = 0; i < L.nstations; ++i) inside += r.QN(i, j);
7219 tput_srv += r.TN(rs - 1, j);
7220 }
7221 const T deficit = T(num_traits<T>::from_double(pop) - inside);
7222 if (!(deficit > Tzero()) || !(tput_srv > Tzero())) return Tzero();
7223 return T(deficit / tput_srv);
7224 }
7225
7226 /** Port of updateRoutingProbabilities: entry selection by throughput ratio. */
7227 void update_routing_probabilities(int) {
7228 for (std::size_t u = 0; u < unique_route_idx.size(); ++u) {
7229 // the reference always takes the reversed order here: its `mod(it,0)`
7230 // guard is NaN, which MATLAB reads as false
7231 const std::size_t idx = unique_route_idx[unique_route_idx.size() - 1 - u];
7232 qn::Layer<T>& L = ensemble[std::size_t(idxhash[idx])];
7233 bool updated = false;
7234 for (const RouteRow& r : route_map) {
7235 if (r.idx != idx) continue;
7236 if (idxhash[r.tidx_caller] < 0) continue;
7237 const std::size_t cl = std::size_t(idxhash[r.tidx_caller]);
7238 const qn::Layer<T>& CL = ensemble[cl];
7239 const LayerResult<T>& cr = results.back()[cl];
7240 // the CALLER's own station, and for a call class the station of
7241 // the entry it targets; the two coincide under `srvn`
7242 const std::size_t cs = station_idx_of(CL, r.tidx_caller);
7243 T Xtot = Tzero();
7244 for (std::size_t k = 0; k < CL.nclasses; ++k) Xtot += cr.TN(cs - 1, k);
7245 if (!(Xtot > Tzero())) continue;
7246 T entry_tput = Tzero();
7247 for (const auto& a : CL.attr_calls)
7248 if (a[3] == r.eidx)
7249 entry_tput += cr.TN(station_idx_of(CL, lqn.parent[a[3]]) - 1, a[0] - 1);
7250 L.set_route(r.cfrom, r.cto, r.nodefrom, r.nodeto, T(entry_tput / Xtot));
7251 updated = true;
7252 }
7253 if (updated) L.refresh_chains();
7254 }
7255 }
7256
7257 // -----------------------------------------------------------------------
7258 // getTranAvg: getTranAvgDecoupled and getTranAvgCoupled
7259 // -----------------------------------------------------------------------
7260
7261 /**
7262 * What a layered transient needs before it can mean anything.
7263 *
7264 * A FLUID ENSEMBLE IS REQUIRED, and not as an implementation shortcut: the
7265 * reference reaches the layer transient through `self.solvers{e}.getTranAvg`,
7266 * which only a transient-capable layer solver has, and the coupled mode
7267 * injects its time-varying demands through the fluid rate schedule
7268 * specifically. An MVA layer has no trajectory to report and no place to
7269 * receive an injection, so an ensemble built on one is refused here rather
7270 * than silently answered with its steady state repeated over a grid.
7271 */
7272 void require_transient_ready() const {
7273 if (opt.layer_solver != "fluid")
7274 throw UnsupportedError(
7275 "SolverLN: the layered transient is the transient OF EACH LAYER, which only the "
7276 "fluid layer solver has; rebuild the ensemble with layer_solver 'fluid'");
7277 }
7278
7279 /** True when the caller named a horizon; otherwise each layer picks its own. */
7280 bool has_finite_horizon() const {
7281 return std::isfinite(opt.timespan_end) && opt.timespan_end > 0.0;
7282 }
7283
7284 /** One layer's trajectory, resampled onto the shared grid. */
7285 struct TranTraj {
7286 std::vector<std::vector<std::vector<double>>> Q, U, Tp, R; ///< [station][class][point]
7287 };
7288
7289 /**
7290 * Run one layer's transient, optionally with an injected rate schedule, and
7291 * return both the reportable block and the trajectory the relaxation reads.
7292 */
7293 void run_layer_transient(std::size_t e, const std::vector<FluidRateSched>& sched,
7294 LnTranLayer& block, TranTraj& traj) {
7295 qn::Layer<T>& L = ensemble[e];
7296 L.refresh_rt();
7297 fluid::FluidOptions fo = opt.layer_fluid;
7298 fo.rate_sched = sched;
7299 // THE WARM START IS REPLAYED HERE, and this is the only place it can be:
7300 // `converged` resets every layer the first time the fixed point settles,
7301 // so a state installed before the solve is gone by now. The steady solve
7302 // ignores an initial state and the transient does not, which is why
7303 // replaying it after the one and before the other loses nothing.
7304 if (e < layer_tran_init.size() && !layer_tran_init[e].empty())
7305 fo.init_sol = layer_tran_init[e];
7306 // No horizon named: this layer picks its own, by the analyzer's own rule
7307 // (thirty mean events of its slowest transition). That is what the
7308 // reference's decoupled path does -- each layer's SolverFluid chooses --
7309 // and it is why an unset timespan is a valid request rather than an error.
7310 const double t_end =
7311 has_finite_horizon() ? opt.timespan_end : fluid::fluid_default_horizon(L, fo);
7312 const std::vector<fluid::FluidTranPoint> pts =
7313 detail::ln_fluid_transient(L, fo, t_end, opt.tran_points, opt.tran_grid);
7314 const std::size_t M = L.nstations, K = L.nclasses, P = pts.size();
7315 block.t.resize(P);
7316 auto alloc = [&](std::vector<std::vector<std::vector<double>>>& A) {
7317 A.assign(M, std::vector<std::vector<double>>(K, std::vector<double>(P, 0.0)));
7318 };
7319 alloc(block.QN);
7320 alloc(block.UN);
7321 alloc(block.TN);
7322 alloc(traj.Q);
7323 alloc(traj.U);
7324 alloc(traj.Tp);
7325 alloc(traj.R);
7326 for (std::size_t p = 0; p < P; ++p) {
7327 block.t[p] = pts[p].t;
7328 for (std::size_t i = 0; i < M; ++i)
7329 for (std::size_t r = 0; r < K; ++r) {
7330 const double q = pts[p].QN(i, r), u = pts[p].UN(i, r), x = pts[p].TN(i, r);
7331 block.QN[i][r][p] = q;
7332 block.UN[i][r][p] = u;
7333 block.TN[i][r][p] = x;
7334 traj.Q[i][r][p] = q;
7335 traj.U[i][r][p] = u;
7336 traj.Tp[i][r][p] = x;
7337 // Residence by Little, which is how the relaxation reads a
7338 // callee's response time off a layer trajectory.
7339 traj.R[i][r][p] = q / std::max(x, GlobalConstants::FineTol);
7340 }
7341 }
7342 }
7343
7344 /**
7345 * Port of getTranAvgDecoupled: freeze the inter-layer demands at the
7346 * converged fixed point and run each layer's transient in isolation.
7347 */
7348 LnTranSolution tran_avg_decoupled() {
7349 require_transient_ready();
7350 if (results.empty()) iterate();
7351 LnTranSolution out;
7352 out.mode = "decoupled";
7353 out.layers.resize(ensemble.size());
7354 for (std::size_t e = 0; e < ensemble.size(); ++e) {
7355 TranTraj tj;
7356 run_layer_transient(e, std::vector<FluidRateSched>(), out.layers[e], tj);
7357 }
7358 return out;
7359 }
7360
7361 /**
7362 * Port of getTranAvgCoupled: reconcile the per-layer transients by waveform
7363 * relaxation, so the layer populations and the inter-layer demands co-evolve
7364 * in model time.
7365 *
7366 * Each layer's fluid transient is driven by TIME-VARYING inter-layer demand
7367 * trajectories taken from the other layers' latest transients, and the loop
7368 * repeats until the trajectories stop moving in sup-norm. Iteration 0 uses
7369 * the frozen equilibrium demands, so it reproduces the decoupled answer
7370 * exactly; at convergence every layer relaxes to its own fixed point, so the
7371 * endpoint equals getAvg.
7372 *
7373 * TWO CHANNELS ARE COUPLED, the task think times (the client delay) and the
7374 * synchronous-call service demands (the caller's client station). Both are
7375 * dominant inter-layer couplings; the intra-layer host service stays at its
7376 * equilibrium value, as in the reference.
7377 */
7378 LnTranSolution tran_avg_coupled() {
7379 require_transient_ready();
7380 // Waveform relaxation reconciles the layers on ONE shared grid, so it
7381 // needs a horizon the layers agree on. With none named there is nothing
7382 // to co-evolve over, and the reference defers to the decoupled path
7383 // rather than inventing one; so does this.
7384 if (!has_finite_horizon()) return tran_avg_decoupled();
7385 if (results.empty()) iterate();
7386 const std::size_t E = ensemble.size();
7387
7388 LnTranSolution out;
7389 out.mode = "coupled";
7390 out.layers.resize(E);
7391 std::vector<TranTraj> traj(E), prev(E);
7392 for (std::size_t e = 0; e < E; ++e)
7393 run_layer_transient(e, std::vector<FluidRateSched>(), out.layers[e], traj[e]);
7394 const std::vector<double> tgrid = out.layers.empty() ? std::vector<double>() : out.layers[0].t;
7395
7396 for (long iter = 1; iter <= opt.ln_transient_iter_max; ++iter) {
7397 prev = traj;
7398 const std::vector<std::vector<FluidRateSched>> sched =
7399 build_rate_sched(recompute_demand(traj, tgrid), tgrid);
7400 for (std::size_t e = 0; e < E; ++e)
7401 run_layer_transient(e, sched[e], out.layers[e], traj[e]);
7402 double gap = 0.0;
7403 for (std::size_t e = 0; e < E; ++e)
7404 for (std::size_t i = 0; i < traj[e].Q.size(); ++i)
7405 for (std::size_t r = 0; r < traj[e].Q[i].size(); ++r)
7406 for (std::size_t p = 0; p < traj[e].Q[i][r].size(); ++p)
7407 gap = std::max(gap, std::fabs(traj[e].Q[i][r][p] - prev[e].Q[i][r][p]));
7408 out.iterations = iter;
7409 out.gap = gap;
7410 if (gap < opt.ln_transient_tol) break;
7411 }
7412 return out;
7413 }
7414
7415 /** The two coupled demand channels, as trajectories on the shared grid. */
7416 struct TranDemand {
7417 std::map<std::size_t, std::vector<double>> thinkt; ///< by task element index
7418 std::map<std::size_t, std::vector<double>> callservt; ///< by call index
7419 };
7420
7421 /**
7422 * Recompute the inter-layer demands pointwise in t, mirroring the SCALAR
7423 * updateThinkTimes and updateMetricsDefault formulas.
7424 *
7425 * They are the same formulas, evaluated at each point of the grid instead of
7426 * at the fixed point: that is what makes the endpoint of the relaxation the
7427 * steady-state answer rather than something near it.
7428 */
7429 TranDemand recompute_demand(const std::vector<TranTraj>& traj,
7430 const std::vector<double>& tgrid) const {
7431 TranDemand out;
7432 const std::size_t ng = tgrid.size();
7433
7434 for (std::size_t t = 1; t <= lqn.ntasks; ++t) {
7435 const std::size_t tidx = lqn.tshift + t;
7436 if (idxhash[tidx] < 0 || lqn.isref[tidx]) continue;
7437 const std::size_t e = std::size_t(idxhash[tidx]);
7438 const qn::Layer<T>& L = ensemble[e];
7439 const std::size_t s = station_idx_of(L, tidx) - 1;
7440 double nj = 0.0;
7441 for (std::size_t c = 1; c <= NT(); ++c) nj = std::max(nj, njobs(tidx, c));
7442 // same closure as update_think_times, so the same gate
7443 const double userthink = dbl(ref_think_time(tidx));
7444 std::vector<double> tk(ng, 0.0);
7445 for (std::size_t p = 0; p < ng; ++p) {
7446 double U = 0.0, X = 0.0;
7447 for (std::size_t r = 0; r < L.nclasses; ++r) {
7448 U += traj[e].U[s][r][p];
7449 X += traj[e].Tp[s][r][p];
7450 }
7451 const double Xs = std::max(X, GlobalConstants::FineTol);
7452 double v = lqn.sched[tidx] == SchedStrategy::INF
7453 ? (nj - U) / Xs - userthink
7454 : nj * std::fabs(1.0 - U) / Xs - userthink;
7456 tk[p] = v + userthink; // total mean, user think included
7457 }
7458 out.thinkt[tidx] = tk;
7459 }
7460
7461 for (std::size_t cidx = 1; cidx <= lqn.ncalls; ++cidx) {
7462 if (lqn.calltype[cidx] != CallType::SYNC) continue;
7463 const std::size_t eidx = lqn.callpair_dst[cidx];
7464 const std::size_t tidx = lqn.parent[eidx];
7465 if (tidx > NT() || idxhash[tidx] < 0) continue;
7466 const std::size_t e = std::size_t(idxhash[tidx]);
7467 const qn::Layer<T>& L = ensemble[e];
7468 const std::size_t s = station_idx_of(L, tidx) - 1;
7469 std::vector<double> Rc(ng, 0.0);
7470 bool any = false;
7471 for (std::size_t r = 0; r < L.nclasses; ++r) {
7472 if (L.classes[r].attr_kind != int(LqnElement::ENTRY)) continue;
7473 if (L.classes[r].attr_idx != eidx) continue;
7474 any = true;
7475 for (std::size_t p = 0; p < ng; ++p) Rc[p] += traj[e].R[s][r][p];
7476 }
7477 if (!any) {
7478 // No entry class of its own in this layer: the entry's work is
7479 // carried by its activities, so their residence is the answer.
7480 for (std::size_t r = 0; r < L.nclasses; ++r)
7481 for (std::size_t p = 0; p < ng; ++p) Rc[p] += traj[e].R[s][r][p];
7482 }
7483 const double cm = dbl(lqn.callproc_mean[cidx]);
7484 for (std::size_t p = 0; p < ng; ++p) Rc[p] *= cm;
7485 out.callservt[cidx] = Rc;
7486 }
7487 return out;
7488 }
7489
7490 /**
7491 * Map the demand trajectories onto per-layer rate schedules, through the
7492 * SAME update maps that place the scalar setService calls.
7493 *
7494 * The injected schedule MODULATES the layer's equilibrium rate by the ratio
7495 * of the transient demand to its steady-state value: passing rate = 1/d(t)
7496 * with nominal = 1/d(end) makes solver_fluid_ratemult's multiplier
7497 * d(end)/d(t), which is exactly 1 at the horizon end. That is what makes the
7498 * layer relax to its UNMODIFIED fixed point regardless of any small mismatch
7499 * between the fluid residence Q/T and the scalar equilibrium demand.
7500 */
7501 std::vector<std::vector<FluidRateSched>> build_rate_sched(
7502 const TranDemand& demand, const std::vector<double>& tgrid) const {
7503 std::vector<std::vector<FluidRateSched>> out(ensemble.size());
7504 const bool want_think =
7505 opt.ln_transient_channels == "both" || opt.ln_transient_channels == "thinkt";
7506 const bool want_call =
7507 opt.ln_transient_channels == "both" || opt.ln_transient_channels == "callservt";
7508
7509 auto add = [&](std::size_t e, std::size_t station, std::size_t cls,
7510 const std::vector<double>& d) {
7511 if (d.empty()) return;
7512 const double dend = d.back();
7513 if (!(dend > GlobalConstants::FineTol)) return; // degenerate steady state
7514 // Bound the transient demand to a physical band around its
7515 // steady-state value: an early-transient throughput near zero sends
7516 // the reciprocal to infinity and the integrator with it.
7517 const double cap = 20.0;
7519 s.station = station;
7520 s.cls = cls;
7521 s.tgrid = tgrid;
7522 s.rates.resize(d.size());
7523 for (std::size_t p = 0; p < d.size(); ++p) {
7524 const double v = std::min(std::max(d[p], dend / cap), dend * cap);
7525 s.rates[p] = 1.0 / v;
7526 }
7527 s.nominal = 1.0 / dend;
7528 out[e].push_back(s);
7529 };
7530
7531 if (want_think)
7532 for (const UpdRow& row : thinkt_map) {
7533 if (idxhash[row.idx] < 0) continue;
7534 const std::size_t e = std::size_t(idxhash[row.idx]);
7535 if (row.node != ensemble[e].clientIdx) continue;
7536 if (lqn.type[row.aidx] != LqnElement::TASK) continue;
7537 if (lqn.sched[row.aidx] == SchedStrategy::REF) continue;
7538 const auto it = demand.thinkt.find(row.aidx);
7539 if (it == demand.thinkt.end()) continue;
7540 add(e, row.node, row.cls, it->second);
7541 }
7542 if (want_call)
7543 for (const UpdRow& row : call_map) {
7544 if (idxhash[row.idx] < 0) continue;
7545 const std::size_t e = std::size_t(idxhash[row.idx]);
7546 if (row.node != ensemble[e].clientIdx) continue;
7547 const auto it = demand.callservt.find(row.aidx);
7548 if (it == demand.callservt.end()) continue;
7549 add(e, row.node, row.cls, it->second);
7550 }
7551 return out;
7552 }
7553
7554 // -----------------------------------------------------------------------
7555 // getEnsembleAvg
7556 // -----------------------------------------------------------------------
7557 LnSolution<T> aggregate() {
7558 const std::size_t N = lqn.nidx;
7559 LnSolution<T> s;
7560 auto mk = [&](std::vector<T>& v, std::vector<bool>& d) {
7561 v.assign(N + 1, Tzero());
7562 d.assign(N + 1, false);
7563 };
7564 std::vector<T> QN, UN, RN, TN, PN, SN, WN, AN;
7565 std::vector<bool> dQ, dU, dR, dT, dP, dS, dW, dA;
7566 mk(QN, dQ); mk(UN, dU); mk(RN, dR); mk(TN, dT);
7567 mk(PN, dP); mk(SN, dS); mk(WN, dW); mk(AN, dA);
7568 std::vector<bool> wn_done(N + 1, false);
7569
7570 for (std::size_t e = 0; e < ensemble.size(); ++e) {
7571 const qn::Layer<T>& L = ensemble[e];
7572 const LayerResult<T>& r = results.back()[e];
7573 const std::size_t clientIdx = L.clientIdx;
7574 // The processors this layer serves: one under `srvn` (and only when
7575 // it is a host layer), every one of them under `flat.cs`. Each is
7576 // charged with the activities that actually run on it, which is
7577 // vacuous under `srvn` because a host layer holds no others.
7578 const bool has_host_server = !L.host_stations.empty();
7579 for (std::size_t hs : L.host_stations) {
7580 const std::size_t hidx = L.stations[hs - 1].attr_idx;
7581 dT[hidx] = true;
7582 TN[hidx] = Tzero();
7583 dP[hidx] = true;
7584 PN[hidx] = Tzero();
7585 for (std::size_t k = 0; k < L.nclasses; ++k) {
7586 if (L.classes[k].completes) {
7587 T t = clientIdx > 0 ? r.TN(clientIdx - 1, k) : Tzero();
7588 const T ts = r.TN(hs - 1, k);
7589 TN[hidx] = T(TN[hidx] + (t > ts ? t : ts));
7590 }
7591 if (L.classes[k].attr_kind == int(LqnElement::ACTIVITY)) {
7592 // the activity does not run on this processor
7593 if (station_idx_of_class(L, k) != hs) continue;
7594 const std::size_t aidx = L.classes[k].attr_idx;
7595 const std::size_t tidx = lqn.parent[aidx];
7596 dP[aidx] = true;
7597 dP[tidx] = true;
7598 PN[aidx] = T(PN[aidx] + r.UN(hs - 1, k));
7599 PN[tidx] = T(PN[tidx] + r.UN(hs - 1, k));
7600 PN[hidx] = T(PN[hidx] + r.UN(hs - 1, k));
7601 }
7602 }
7603 dT[hidx] = false; // NaN in the reference, for consistency with LQNS
7604 }
7605
7606 for (std::size_t k = 0; k < L.nclasses; ++k) {
7607 const int kind = L.classes[k].attr_kind;
7608 // the station this class is actually served at, which under
7609 // `flat.cs` is the processor of an activity or the called task
7610 // of a call rather than the one server the layer used to have
7611 const std::size_t serverIdx = station_idx_of_class(L, k);
7612 if (kind == int(LqnElement::TASK)) {
7613 const std::size_t tidx = L.classes[k].attr_idx;
7614 if (has_host_server && !dT[tidx]) {
7615 dT[tidx] = true;
7616 TN[tidx] = r.TN(clientIdx - 1, k);
7617 }
7618 } else if (kind == int(LqnElement::ENTRY)) {
7619 const std::size_t eidx = L.classes[k].attr_idx;
7620 dS[eidx] = true;
7621 // getEnsembleAvg.m:84-92: a phase-2 entry reports the
7622 // CALLER's view, which is what residt now holds.
7623 SN[eidx] = (has_phase2 && dbl(servt_ph2[eidx]) > GlobalConstants::FineTol)
7624 ? residt[eidx]
7625 : servt[eidx];
7626 if (has_host_server && !dT[eidx]) {
7627 dT[eidx] = true;
7628 TN[eidx] = r.TN(clientIdx - 1, k);
7629 }
7630 } else if (kind == int(LqnElement::CALL)) {
7631 const std::size_t cidx = L.classes[k].attr_idx;
7632 const std::size_t aidx = lqn.callpair_src[cidx];
7633 if (lqn.calltype[cidx] == CallType::SYNC) {
7634 dS[aidx] = true;
7635 SN[aidx] = T(SN[aidx] + r.RN(serverIdx - 1, k) * lqn.callproc_mean[cidx]);
7636 }
7637 dQ[aidx] = true;
7638 QN[aidx] = T(QN[aidx] + r.QN(serverIdx - 1, k));
7639 } else if (kind == int(LqnElement::ACTIVITY)) {
7640 const std::size_t aidx = L.classes[k].attr_idx;
7641 const std::size_t tidx = lqn.parent[aidx];
7642 dQ[tidx] = true;
7643 QN[tidx] = T(QN[tidx] + r.QN(serverIdx - 1, k));
7644 dT[aidx] = true;
7645 dQ[aidx] = true;
7646 TN[aidx] = T(TN[aidx] + r.TN(serverIdx - 1, k));
7647 dS[aidx] = true;
7648 SN[aidx] = T(SN[aidx] + r.RN(serverIdx - 1, k));
7649 dR[aidx] = true;
7650 RN[aidx] = T(RN[aidx] + r.RN(serverIdx - 1, k));
7651 dW[aidx] = true;
7652 dW[tidx] = true;
7653 WN[aidx] = residt[aidx];
7654 if (!wn_done[aidx]) {
7655 WN[tidx] = T(WN[tidx] + residt[aidx]);
7656 wn_done[aidx] = true;
7657 }
7658 QN[aidx] = T(QN[aidx] + r.QN(serverIdx - 1, k));
7659 }
7660 }
7661 }
7662
7663 for (std::size_t e = 1; e <= lqn.nentries; ++e) {
7664 const std::size_t eidx = lqn.eshift + e;
7665 const std::size_t tidx = lqn.parent[eidx];
7666 dU[tidx] = true;
7667 dU[eidx] = true;
7668 // getEnsembleAvg.m:186-198: the server is busy through BOTH phases,
7669 // so the utilization is the full service time and not SN, which for
7670 // a phase-2 entry has been cut down to the caller's view above.
7671 UN[eidx] = (has_phase2 && dbl(servt_ph2[eidx]) > GlobalConstants::FineTol)
7672 ? T(TN[eidx] * (servt_ph1[eidx] + servt_ph2[eidx]))
7673 : T(TN[eidx] * SN[eidx]);
7674 T ps = Tzero();
7675 bool any = false;
7676 for (std::size_t a : lqn.actsof[eidx])
7677 if (dP[a]) {
7678 ps += PN[a];
7679 any = true;
7680 }
7681 if (!lqn.actsof[eidx].empty()) {
7682 dP[eidx] = true;
7683 PN[eidx] = any ? ps : Tzero();
7684 }
7685 for (std::size_t a : lqn.actsof[tidx]) {
7686 dU[a] = true;
7687 UN[a] = T(TN[a] * SN[a]);
7688 }
7689 UN[tidx] = T(UN[tidx] + UN[eidx]);
7690 }
7691
7692 // AN IGNORED ELEMENT IS IDLE, NOT UNDEFINED, and the two are different cells.
7693 // Its component holds no reference task, so nothing reaches it and every
7694 // measure it HAS is zero -- but the measures its kind never has stay
7695 // undefined, exactly as they do for a reachable element. A flat zero over
7696 // all six columns broke the table's NaN mask (a processor with a queue
7697 // length of 0, an arrival rate reported where no solver reports one), and
7698 // the mask is part of the answer: see _kb/06-solver-catalog.md. The
7699 // relabelling below reports Q from U, U from P and R from S, so QN and RN
7700 // are dead here and are not written.
7701 for (std::size_t i = 1; i <= N; ++i)
7702 if (ignore[i]) {
7703 PN[i] = Tzero(); // every kind reports a utilization
7704 dP[i] = true;
7705 dA[i] = false; // nothing reports an arrival rate on an LQN
7706 const bool host = lqn.type[i] == LqnElement::HOST;
7707 const bool task = lqn.type[i] == LqnElement::TASK;
7708 const bool entry = lqn.type[i] == LqnElement::ENTRY;
7709 UN[i] = Tzero();
7710 dU[i] = !host;
7711 SN[i] = Tzero();
7712 dS[i] = !host && !task;
7713 WN[i] = Tzero();
7714 dW[i] = !host && !entry;
7715 TN[i] = Tzero();
7716 dT[i] = !host;
7717 }
7718
7719 // the reference's final relabelling: Q <- U, U <- P, R <- S
7720 s.QN = UN; s.defined_Q = dU;
7721 s.UN = PN; s.defined_U = dP;
7722 s.RN = SN; s.defined_R = dS;
7723 s.TN = TN; s.defined_T = dT;
7724 s.AN = AN; s.defined_A = dA;
7725 s.WN = WN; s.defined_W = dW;
7726 s.iterations = iterations_done;
7727 s.converged = did_converge;
7728 return s;
7729 }
7730};
7731
7732} // namespace ln
7733} // namespace line
7734
7735// Deliberately at the FOOT of the file: lqn_analyzers.h needs LayerResult and
7736// SolverLN complete, and this file needs its lqn_overtake_prob_markov, so the
7737// two are mutually dependent. Either include order works -- whichever header a
7738// translation unit names first, the other is fully parsed before the templates
7739// above are instantiated -- and the forward declaration near the top of this
7740// file is what makes the call inside update_metrics resolve.
7742
7743#endif // LINE_SOLVERS_LN_SOLVER_LN_H
Convolution of a sequence of matrix-exponential laws.
Minimal-order acyclic phase-type fit of the first three moments (matlab/lib/kpctoolbox/aph/aph_fit....
Composition of two matrix-exponential distributions given in (alpha, T) form.
Malformed or inconsistent input (dimensions, negative populations, ...).
Definition error.h:37
InputError(const std::string &what)
Definition error.h:39
std::size_t cols() const
Definition matrix.h:90
std::size_t rows() const
Definition matrix.h:89
bool empty() const
Definition matrix.h:92
UnsupportedError(const std::string &what)
Definition error.h:51
Convergence controller for an ensemble whose layers are solved by a NOISY method (simulation,...
const std::vector< double > & state_refpath_stages() const
Per reference-path stage key, the stage mean the last iteration installed.
Definition solver_ln.h:794
const std::string & state_interlock_method() const
The interlock tracking method this run USES: "ilrate", "refpath" or "none".
Definition solver_ln.h:792
LnSolution< T > box_bounds()
Majumdar-Woodside robust box bounds, reported in the shape of a solution.
Definition solver_ln.h:655
LnLayerBlocks layer_blocks() const
Port of LayeredNetwork.layerBlocks: the block-diagonal layout of the layers in the aggregate (station...
Definition solver_ln.h:699
const std::vector< std::vector< LayerResult< T > > > & iteration_results() const
Diagnostic access to the per-iteration layer results, for the regression.
Definition solver_ln.h:774
LnInterlockChoice interlock_method_for() const
Port of SolverLN.interlockMethodFor: the method a solve would use, asked through the same query build...
Definition solver_ln.h:802
const std::vector< T > & state_util() const
Definition solver_ln.h:784
SolverLN(const LqnStruct< T > &lqn_in, const LnOptions &options)
Definition solver_ln.h:506
LnSensTable< T > get_sensitivity_table(const sens::SensOptions &sopt)
Port of @SolverLN/getSensitivityTable: solve the ensemble, then concatenate each layer solver's own t...
Definition solver_ln.h:610
LnTranSolution get_tran_avg()
Port of @SolverLN/getTranAvg: the block-diagonal aggregate transient.
Definition solver_ln.h:591
const std::vector< T > & state_thinkt() const
Definition solver_ln.h:778
std::size_t nlayers() const
Definition solver_ln.h:688
const std::vector< Distrib< T > > & state_callservtproc() const
Definition solver_ln.h:783
std::vector< LnCdf > get_cdf_respt()
Port of @SolverLN/getCdfRespT: the per-entry response-time distribution.
Definition solver_ln.h:552
const std::vector< T > & state_callresidt() const
Definition solver_ln.h:780
const std::string & state_lnmethod() const
The method the layers were built for: "srvn.ph", "srvn.cs" or "moment3".
Definition solver_ln.h:786
const std::vector< T > & state_tput() const
Definition solver_ln.h:777
void set_tran_grid(const std::vector< double > &g)
Install the explicit output grid of the layered transient; see LnOptions::tran_grid.
Definition solver_ln.h:771
const std::vector< T > & state_servt() const
Definition solver_ln.h:775
static std::vector< std::string > list_valid_methods()
Port of SolverLN.listValidMethods.
Definition solver_ln.h:502
const std::vector< Distrib< T > > & state_servtproc() const
Definition solver_ln.h:781
const std::vector< Distrib< T > > & state_thinktproc() const
Definition solver_ln.h:782
LnSolution< T > get_ensemble_avg()
Port of getEnsembleAvg: run the iteration and aggregate onto LQN elements.
Definition solver_ln.h:528
const std::vector< T > & state_callservt() const
Definition solver_ln.h:779
const std::vector< qn::Layer< T > > & layers() const
Definition solver_ln.h:689
const std::vector< T > & state_residt() const
Definition solver_ln.h:776
void init_from_marginal(const Matrix< double > &n)
Port of LayeredNetwork.initFromMarginal: split an aggregate (M x K) mean queue-length matrix into per...
Definition solver_ln.h:739
A layer network: everything a NetworkStruct holds, plus the LQN back-mapping.
Definition qn_layer.h:52
A network plus its refreshed NetworkStruct.
std::vector< std::vector< bool > > disabled
std::vector< Station< T > > stations
stations[k-1] is the k-th station
static void iter(long k, const char *fmt,...)
Report iteration k of the current loop.
static void step(const char *fmt,...)
Write one progress line.
static bool is_acyclic_generator(const Matrix< T > &S)
True when the phase graph of S has no cycle.
Definition workflow.h:670
static PhLaw< T > compose_loop_geometric(const PhLaw< T > &body, const T &count)
Geometric repetition of a phase-type law, the POST_LOOP semantics.
Definition workflow.h:598
static PhLaw< T > compose_serial(const PhLaw< T > &a, const PhLaw< T > &b)
Serial composition: the second law starts when the first absorbs.
Definition workflow.h:476
static PhLaw< T > compose_mixture(const std::vector< PhLaw< T > > &laws, const std::vector< T > &probs)
Probabilistic mixture: a block-diagonal generator whose initial vector picks branch i with probabilit...
Definition workflow.h:561
What refreshProcessRepresentations and refreshLST compute FROM a distribution: the (D0,...
The exception types the port throws.
Matrix exponential by scaling and squaring with a diagonal Pade approximant.
Activities belonging to each branch of an AND-join.
The fork-join fixed point that drives one inner MVA solve.
The fork-join transform SolverMVA applies before solving a layer that contains a Fork.
Mean and variance of a k-of-n (quorum) join completion time, from the mean and variance of each branc...
Response-time distribution by tagged fluid: a port of solver_fluid_passage_time.m,...
The fluid solver's outermost entry point: @@SolverFLD/runAnalyzer.m's method resolution over solver_f...
Dense linear algebra over the templated number type: products, identity, inverse, and powers.
Running progress log of a LINE solver run (the "solver console").
The @SolverLN methods that solver_ln.h does not carry.
Majumdar-Woodside robust box bounds on the throughput of a layered network.
Standalone LQN routines that SolverLN needs but does not contain.
Phase-type composition of an LQN activity graph, the machinery behind SolverLN method 'srvn....
Synchronous call DAG carrying reference-task customers into a layer.
LayeredNetworkStruct, the flattened description of a layered queueing network.
Port of solver_mam_analyzer.m: one inner solve, choosing the analyzer that fits the model and the req...
Dense matrix and non-owning view.
Port of @@SolverMVA/mvaDispatch.m: one inner solve, choosing the analyzer that fits the model.
EntryWorkflow< T > entry_workflow(const ::line::lqn::LqnStruct< T > &lqn, std::size_t eidx, bool with_calls)
Activity graph of LQN entry EIDX as a Workflow.
Definition lqn_ph.h:104
PhLaw< T > ph_law_of(const Distrib< T > &d)
The (alpha, S) pair of a phase-type Distrib, the form the composition rules take.
Definition lqn_ph.h:52
PhLaw< T > serial_law(Workflow< T > &wf)
Composed law of a workflow in which the branches of an AND fork are SERIAL rather than concurrent,...
Definition lqn_ph.h:241
std::pair< T, T > ph_moments(const std::vector< T > &alpha, const Matrix< T > &S)
First two moments of a phase-type law without building a Distrib, which is what the layered fixed poi...
Definition lqn_ph.h:258
void sn_fj_nodevisits_mmt(qn::NetworkStruct< T > &sn)
Rewrite sn.nodevisits with the MMT correction.
T sn_compat_scaling(const Matrix< T > &compat, const std::vector< double > &counts, const std::vector< T > &rates, const std::vector< T > &n)
Rate scaling eta(n) a compatibility declaration imposes on its station.
mva::AvgResult< T > solver_ctmc_run_analyzer_any(const NetworkStruct< T > &sn, const CtmcOptions &opt)
Solve on whichever path applies and format, mirroring solver_ctmc_run_analyzer.
std::vector< std::vector< std::size_t > > fj_branch_members(const LqnBranchView< T > &lqn, std::size_t joinaidx)
Branch membership of an AND-join.
FJQuorumMomentsResult< T > fj_quorum_moments(const std::vector< T > &branchMeans, const std::vector< T > &branchVars, std::size_t k)
Mean and variance of a k-of-n (quorum) join completion time, from the mean and variance of each branc...
double fluid_default_horizon(const qn::NetworkStruct< T > &sn, const FluidOptions &opt)
The horizon a transient runs to when the caller gives none.
FluidLayout fluid_layout(const qn::NetworkStruct< T > &sn)
Port of the layout half of solver_fluid_odes.m.
Definition fluid_odes.h:282
std::vector< FluidTranPoint > solver_fluid_transient(const qn::NetworkStruct< T > &sn, const FluidOptions &opt, double t_end, std::size_t points=101, const std::vector< double > &out_grid=std::vector< double >())
Port of @@SolverFLD/getTranAvg: the metrics along the trajectory, not just at the fixed point.
FluidSolution solver_fluid(const qn::NetworkStruct< T > &sn_in, const FluidOptions &opt, qn::NetworkStruct< T > *sn_out=nullptr)
Port of solver_fluid_analyzer.m: dispatch on the method, refit the non-exponential FCFS stations the ...
FluidPassage fluid_passage_time(const qn::NetworkStruct< T > &sn, const std::vector< double > &x_steady, std::size_t ist, std::size_t cls, double tol=1e-4, std::size_t points=201, const FluidClosure &closure=FluidClosure())
Response-time CDF at station ist for class cls, both 1-based.
FluidSolution solver_fluid_run_analyzer(const qn::NetworkStruct< T > &sn, const FluidOptions &opt, qn::NetworkStruct< T > *sn_out=nullptr, qn::NetworkStruct< T > *refreshed_out=nullptr, solvers::CacheMetrics< T > *cache_out=nullptr)
Port of @@SolverFLD/runAnalyzer.m: resolve the method, route to the function the reference routes to,...
SchedStrategy
Scheduling disciplines, with the values of MATLAB SchedStrategy.
Definition lang_types.h:181
PrecedenceType
Activity precedence kinds, with the values of MATLAB ActivityPrecedenceType.
Definition lang_types.h:472
LqnElement
LQN element kinds, with the values of MATLAB LayeredNetworkElement.
Definition lang_types.h:466
RoutingStrategy
Routing strategies, with the values of MATLAB RoutingStrategy.
Definition lang_types.h:391
CallType
Call kinds, with the values of MATLAB CallType.
Definition lang_types.h:469
Distrib< T > aph_fit_mean_scv(const T &mean, const T &scv)
APH.fitMeanAndSCV(MEAN, SCV), through mam::aph_fit_mean_scv.
JobClassType
Job class kinds, with the values of MATLAB JobClassType.
Definition lang_types.h:369
std::function< std::vector< T >(const std::vector< T > &)> CdScaling
A class-dependent scaling map, sn.cdscaling.
Definition lang_types.h:731
NodeType
Node kinds, with the values of MATLAB NodeType.
Definition lang_types.h:326
T dist_moment(const Distrib< T > &d, unsigned k)
The k-th raw moment.
void lqn_fwd_rendezvous(LqnStruct< T > &lqn)
Replace every forwarding chain reachable from a synchronous call by caller-side pseudo rendezvous cal...
Definition lqn_helpers.h:77
T lqn_overtake_prob_markov(const LqnStruct< T > &lqn, const std::vector< T > &servt, const std::vector< T > &callresidt, const std::vector< T > &tput, std::size_t eidx, const T &xj)
Overtaking probability at a server entry, through the LQNS phased-server chain rather than the reduce...
LnInterlockChoice ln_interlock_method(const LqnStruct< T > &lqn, const std::string &requested, bool interlocking, const std::string &lnmethod)
Port of matlab/src/solvers/LN/ln_interlock_method.m, asked as a query.
fluid::FluidOptions::RateSched FluidRateSched
One (station, class) rate trajectory injected into a layer's closing ODE.
Definition solver_ln.h:147
LqnRefRoutes< T > lqn_ref_routes(const LqnStruct< T > &lqn, const std::vector< std::size_t > &callers, double maxpaths=32.0, const std::vector< std::size_t > &server_set={})
Resolve the reference routes into the layer whose callers are CALLERS.
LqnBoxBounds< T > lqn_boxbounds(const LqnStruct< T > &lqn)
Evaluate the box bounds of lqn.
MamSolution< T > mam_dispatch(const qn::NetworkStruct< T > &L, const MamOptions &opt_in)
The ladder.
AphPair< T > aph_convseq(const std::vector< AphPair< T > > &seq)
Convolve the sequence, i.e.
Definition aph_convseq.h:37
AphPair< T > aph_simplify(const AphPair< T > &d1, const AphPair< T > &d2, const T &p1, const T &p2, AphPattern pattern)
Compose two matrix-exponential laws, as aph_simplify.m does.
AphFitResult< T > aph_fit(const T &e1, const T &e2, const T &e3, unsigned nmax, const T &tol)
Fit an APH(n) with n <= nmax to the raw moments e1, e2, e3.
Definition aph_fit.h:176
FjMmt< T > fj_mmt(const qn::NetworkStruct< T > &L)
Build the transformed layer.
Definition fj_mmt.h:444
bool mva_carries_interlock(const qn::NetworkStruct< T > &L, const MvaOptions &opt)
True when the MVA path this model already dispatches to carries a class-level interlock matrix (Frank...
MvaSolution< T > fj_fixed_point(const qn::NetworkStruct< T > &L, FjMmt< T > &tr, std::vector< T > &lam, const MvaOptions &opt, InnerSolve inner)
Drive the fork-join fixed point of a transformed model to convergence.
Definition fj_driver.h:435
DispatchResult< T > mva_dispatch(const qn::NetworkStruct< T > &L, const MvaOptions &opt, const Matrix< T > &init_sol)
The ladder itself.
NcSolution< T > solver_nc_solve(const qn::NetworkStruct< T > &L_in, const NcSolverOptions &opt_in)
The gates, the multiserver conversion and the dispatch of @@SolverNC/runAnalyzer.m,...
SensTable< T > solver_sensitivity_table(qn::NetworkStruct< T > &sn, const SensOptions &opt, bool exact_available, const std::function< mva::MvaSolution< T >()> &solve)
Build the sensitivity table of sn under solve.
SsaSolution solver_ssa(const qn::NetworkStruct< T > &sn, const SsaOptions &opt, std::vector< SsaCacheRatio > *cache=nullptr)
@@SolverSSA/runAnalyzer itself: the engine the method selects, then the result assembly the reference...
Conservation laws of a layered queueing network, enumerated from its structure.
Definition aoi_dist2ph.h:52
Matrix< T > expm(const Matrix< T > &A)
Matrix exponential exp(A).
Definition expm.h:141
Port of solver_nc_analyzer.m, solver_ncld_analyzer.m and @@SolverNC/ncDispatch.m: one inner solve,...
One SolverLN layer: a NetworkStruct plus the LQN annotations that say which element of the layered mo...
Total service rate of a station served by heterogeneous server pools with a class-compatibility graph...
Post-MMT node visits of a fork-join model.
Port of solver_ctmc_fcr_waitq.m: the reachability-built generator of a model whose finite capacity re...
SolverFluid: the closing method, a port of solver_fluid.m, solver_fluid_iteration....
SolverMVA over a SolverLN layer.
The SolverNC class surface: @@SolverNC/runAnalyzer.m and the gates around it.
Performance sensitivities with respect to service rates.
The SolverSSA entry surface: a port of @@SolverSSA/runAnalyzer.m's method whitelist,...
Where each (station, class) block sits in the state vector.
Definition fluid_odes.h:86
std::vector< std::vector< std::size_t > > qidx
0-based first index of (i,r)
Definition fluid_odes.h:88
std::size_t nstates
length of the state vector
Definition fluid_odes.h:87
std::vector< std::vector< bool > > enabled
whether (i,r) is served at all
Definition fluid_odes.h:90
options.config.rate_sched: explicit per-(station, class) rate trajectories, the third source solver_f...
Controls, defaulting to SolverOptions('Fluid') in the reference.
double tol
absolute and relative tolerance handed to the integrator
What the analyzer returns, in the same shape as the MVA solver's result.
std::vector< double > XN
std::vector< double > xvec
the converged fluid state
std::vector< double > CN
static Distrib exp_rate(const T &r)
Definition lang_types.h:945
static Distrib phase_type(const std::vector< T > &alpha, const Matrix< T > &A, bool acyclic)
PH / APH given by (alpha, A): D0 = A and D1 = (-A e) alpha.
static Distrib disabled_dist()
Definition lang_types.h:988
static T ph_moment(const std::vector< T > &alpha, const Matrix< T > &A, unsigned k)
The k-th raw moment of a phase-type (alpha, A): k!
static Distrib immediate()
The Immediate singleton.
Definition lang_types.h:977
static Distrib exp_mean(const T &m)
Definition lang_types.h:930
static constexpr double Immediate
Rate of an Immediate distribution; its mean is 1/Immediate = 1e-8.
Definition lang_types.h:766
static constexpr double FineTol
Definition lang_types.h:760
static constexpr double Zero
Definition lang_types.h:762
static constexpr double CoarseTol
Definition lang_types.h:761
Per-layer results of one iteration, the [QN,UN,RN,TN,AN,WN] of getAvg.
Definition solver_ln.h:365
A CDF sampled on a grid, the [F, t] pair MATLAB's evalCDF returns.
Definition solver_ln.h:386
bool empty() const
Definition solver_ln.h:388
std::vector< double > t
Definition solver_ln.h:387
std::vector< double > cdf
Definition solver_ln.h:387
The interlock tracking method a run will use, and why it differs from the request.
LayeredNetwork.layerBlocks: where each layer's block sits in the aggregate.
Definition solver_ln.h:419
std::vector< std::size_t > msz
Definition solver_ln.h:420
std::vector< std::size_t > roff
Definition solver_ln.h:420
std::vector< std::size_t > coff
Definition solver_ln.h:420
std::vector< std::size_t > ksz
Definition solver_ln.h:420
Options of SolverLN.
Definition solver_ln.h:266
std::size_t tran_points
Output points per layer trajectory, and the relaxation's own grid size.
Definition solver_ln.h:347
std::vector< double > tran_grid
An EXPLICIT output grid for the layered transient, replacing the uniform tran_points one when it is n...
Definition solver_ln.h:360
std::string relax
Definition solver_ln.h:289
fluid::FluidOptions layer_fluid
Options handed to each layer when layer_solver is fluid.
Definition solver_ln.h:310
std::string layer_solver
Which solver runs each layer: mva, nc, fluid or ssa.
Definition solver_ln.h:308
std::string interlock_method
options.config.interlock_method: how the interlock is TRACKED once interlocking is on.
Definition solver_ln.h:280
mva::MvaOptions layer
Options handed to each layer solver; SolverMVA defaults.
Definition solver_ln.h:292
std::string method
options.method, which selects WHAT is reported and not merely how:
Definition solver_ln.h:326
ssa::SsaOptions layer_ssa
Options handed to each layer when layer_solver is ssa.
Definition solver_ln.h:314
std::string ln_transient_channels
options.config.ln_transient_channels: which inter-layer coupling is injected, both,...
Definition solver_ln.h:343
double interlock_maxpaths
options.config.interlock_maxpaths: refpath refuses a layer carrying more routes into it.
Definition solver_ln.h:282
std::string ln_transient
options.config.ln_transient: how the per-layer transients are coupled.
Definition solver_ln.h:333
nc::NcSolverOptions layer_nc
Options handed to each layer when layer_solver is nc.
Definition solver_ln.h:312
std::string interlock_refpath_scope
options.config.interlock_refpath_scope: merging transforms only the layers where two or more callers ...
Definition solver_ln.h:288
double timespan_end
options.timespan(2): the transient horizon; infinite means none is set.
Definition solver_ln.h:345
long ln_transient_iter_max
options.config.ln_transient_iter_max and ..._tol of the relaxation.
Definition solver_ln.h:335
getSensitivityTable of the ensemble: the layer tables under a Layer column.
Definition solver_ln.h:434
std::vector< std::string > layer_methods
Per layer, the branch that layer took; empty for a layer with no solver.
Definition solver_ln.h:441
std::vector< sens::SensTable< T > > layer_tables
Per layer, the analytic Jacobian where that layer produced one.
Definition solver_ln.h:445
std::vector< Row > rows
Definition solver_ln.h:439
std::string method
The summary label: the common branch, or "mixed" when they differ.
Definition solver_ln.h:443
The LQN-level answer, indexed by element 1..nidx.
Definition solver_ln.h:371
std::vector< T > WN
Definition solver_ln.h:372
std::vector< bool > defined_W
Definition solver_ln.h:373
std::vector< bool > defined_U
Definition solver_ln.h:373
std::vector< T > QN
Definition solver_ln.h:372
std::vector< bool > defined_R
Definition solver_ln.h:373
std::vector< T > RN
Definition solver_ln.h:372
std::vector< T > AN
Definition solver_ln.h:372
std::vector< bool > defined_Q
Definition solver_ln.h:373
bool is_bound
True when the numbers are a BOUND (method = mw.upper / mw.lower) rather than the fixed point.
Definition solver_ln.h:382
std::vector< bool > defined_A
Definition solver_ln.h:373
std::vector< T > UN
Definition solver_ln.h:372
std::vector< bool > defined_T
Definition solver_ln.h:373
std::vector< T > TN
Definition solver_ln.h:372
options.config.stochiter_* of SolverOptions.m, with its defaults.
Definition solver_ln.h:464
double a0
Robbins-Monro step immediately after burn-in.
Definition solver_ln.h:466
long conseq
consecutive sub-tolerance iterations required to stop
Definition solver_ln.h:468
long burnin
Picard iterations before the step decay starts.
Definition solver_ln.h:465
double alpha
step decay exponent, in (0.5, 1]
Definition solver_ln.h:467
double relax_burnin
Relaxation in force during burn-in, i.e.
Definition solver_ln.h:471
One layer's block of the layered transient.
Definition solver_ln.h:399
std::vector< std::vector< std::vector< double > > > QN
[station][class][point]
Definition solver_ln.h:402
std::vector< double > t
the output grid, shared by every series below
Definition solver_ln.h:400
std::vector< std::vector< std::vector< double > > > UN
Definition solver_ln.h:402
std::vector< std::vector< std::vector< double > > > TN
Definition solver_ln.h:402
The layered transient: one block per layer, plus how it was produced.
Definition solver_ln.h:425
std::string mode
"coupled" or "decoupled"
Definition solver_ln.h:427
long iterations
waveform-relaxation sweeps; 0 when decoupled
Definition solver_ln.h:428
double gap
final sup-norm trajectory change, coupled only
Definition solver_ln.h:429
std::vector< LnTranLayer > layers
Definition solver_ln.h:426
std::vector< std::size_t > targets
absolute entry indices, in declaration order
Definition lqn_struct.h:204
std::size_t caller
absolute index of the dispatching activity
Definition lqn_struct.h:202
The options SolverMVA reads.
Definition mva_types.h:31
Class-level results, the [Q,U,R,T,C,X] of the MATLAB analyzers.
Definition mva_types.h:96
std::vector< T > X
Definition mva_types.h:98
std::vector< T > C
Definition mva_types.h:98
Controls, defaulting to SolverOptions('NC') in the reference.
Definition nc_types.h:33
The name-value contract of getSensitivityTable.
bool simulation
True when the callback is a simulator, which widens the default step.
One (station, class) row of the table.
What the table carries, plus the branch that produced it.
std::vector< SensRow< T > > rows
std::string method
"exact" or "fd", the branch actually taken
Controls, defaulting to SolverOptions('SSA') in the reference.
Definition ssa_types.h:69