6. Allocation of Computing Resources#

A QUEENS scheduler splits a fixed core budget with two parameters: how many jobs run at the same time (num_jobs) and how many cores each single job gets (num_procs).

This tutorial runs the same study under different allocations, shows when an allocation leads to idle cores, and fixes that by splitting up the study into multiple optimally sized runs.

Absolute timings depend on the computer used, so every result below is read as a ratio. Those ratios carry over to any hardware; the runtimes themselves do not.

The Parameters#

parameter

meaning

num_jobs

how many jobs run at the same time

num_procs

how many cores each single job may use

cores in use = num_jobs × num_procs

The scheduler passes num_procs to the driver’s run(...); how the driver spends those cores is up to you. Which split is right depends on the solver, e.g., a serial solver wants many one-core jobs.

Cluster schedulers include a third parameter: num_nodes - the number of cluster nodes allocated per job; num_procs then assigns cores per node, so one job gets num_nodes × num_procs cores while num_jobs still sets how many such jobs run in parallel.


Example:

On a cluster with 24-core nodes, set num_jobs = 3, num_nodes = 2, and num_procs = 24 to run three simultaneous jobs, each using 48 cores. Each job spans two nodes (2 × 24 = 48 cores), so the three jobs together occupy num_jobs × num_nodes × num_procs = 144 cores.

[1]:
import logging
import os
import subprocess
import sys
import time
from pathlib import Path

import numpy as np

from queens.distributions import Uniform
from queens.drivers._driver import Driver
from queens.global_settings import GlobalSettings
from queens.iterators import MonteCarlo
from queens.main import run_iterator
from queens.models import Simulation
from queens.parameters import Parameters
from queens.schedulers import Local
from queens.utils.io import load_result

logging.getLogger("distributed").setLevel(logging.WARNING)              # hides unnecessary output
os.environ["DASK_DISTRIBUTED__LOGGING__DISTRIBUTED"] = "warning"        # hides unnecessary output

N_POINTS = 10000000     # points per pi estimate; USER: try different values to feel the difference
SEED = 17               # master seed of the whole tutorial

Example: Estimating \(\pi\)#

Throw points \((x, y)\) at random such that they are uniformly distributed on a unit square \([0,1] \times [0,1]\). The probability that a single point lands inside (hits) the quarter circle is purely geometric - the ratio of the two areas:

\[P(x^2 + y^2 \le 1) = \frac{\text{area of quarter circle}}{\text{area of square}} = \frac{\pi/4}{1} = \frac{\pi}{4}\]

This is not an approximation: \(\pi\) is contained exactly in this probability before a single point is thrown. The probability can then be estimated by repetition - the more points you throw, the closer the fraction of hits gets to \(\pi/4\) (law of large numbers):

\[\hat{p}_N = \frac{\text{hits}}{N} \;\xrightarrow{N \to \infty}\; \frac{\pi}{4} \qquad \Rightarrow \qquad \pi_{\text{est}} = 4 \cdot \hat{p}_N \to \pi\]

Examples:

With 10 throws (\(N = 10\)) and 8 hits in the quarter circle

\[\hat{p}_{10} = \frac{8}{10} = 0.8 \qquad \Rightarrow \qquad \pi_{\text{est}} = 4 \cdot 0.8 = 3.2\]

With \(10^6\) throws and 785398 hits in the quarter circle

\[\hat{p}_{10^6} = \frac{785398}{10^6} = 0.785398 \qquad \Rightarrow \qquad \pi_{\text{est}} = 4 \cdot 0.785398 = 3.141592\]

The error decays with a rate of \(1/\sqrt{N}\), so a decent estimate needs millions of points.

Step 1: One Estimate on One Core#

count_hits() counts the hits inside the quarter circle. N_POINTS is a runtime parameter: the higher this number, the more time the program takes on a single core - that allows the parallelization below to speed up the estimation.

[2]:
def count_hits(n_points, seed):
    rng = np.random.default_rng(seed)
    x = rng.random(n_points)
    y = rng.random(n_points)
    return int(np.count_nonzero(x**2 + y**2 <= 1.0))


start = time.perf_counter()
hits = count_hits(N_POINTS, SEED)
print(f"hits = {hits}")
pi_est = 4.0 * hits / N_POINTS
print(f"pi ~ {pi_est:.5f}  ({time.perf_counter() - start:.2f} sec with 1 core)")
hits = 7854768
pi ~ 3.14191  (0.08 sec with 1 core)

Step 2: One Estimate on Several Cores#

To make a single estimate faster, split it into two parts that will carry through the rest of the tutorial:

  • split_work() - the cheap, serial part: divide N_POINTS into num_chunks chunks, each with its own seed so the chunks draw independent points.

  • run_chunks() - the expensive, parallel part: one worker process per chunk counts its hits, then the hits are combined.

estimate_pi() is just the two parts glued together.

Don’t expect the speedup to match the increase in cores: with num_procs = 4 the ideal would be a 4x speedup, but you will measure clearly less. Every worker is a fresh Python process, and starting one - launching the interpreter, importing numpy, etc. - takes a fixed amount of time that stays the same no matter how many cores you add.

[3]:
WORKER = """
import sys, numpy as np
n_points, seed = int(sys.argv[1]), int(sys.argv[2])
rng = np.random.default_rng(seed)
x = rng.random(n_points)
y = rng.random(n_points)
print(int(np.count_nonzero(x ** 2 + y ** 2 <= 1.0)))
"""


def split_work(n_points, seed, num_chunks):
    """Serial part: split one estimate into `num_chunks` chunks, each with its own seed."""
    seeds = np.random.SeedSequence(seed).generate_state(num_chunks)
    counts = [n_points // num_chunks] * num_chunks
    counts[-1] += n_points - sum(counts)  # last chunk takes the remainder
    return list(zip(counts, seeds))


def run_chunks(chunks):
    """Parallel part: one worker process per chunk, combine the hits."""
    workers = [
        subprocess.Popen(
            [sys.executable, "-c", WORKER, str(n), str(s)], stdout=subprocess.PIPE, text=True
        )
        for n, s in chunks
    ]
    hits = sum(int(w.communicate()[0]) for w in workers)
    total = sum(n for n, _ in chunks)
    return 4.0 * hits / total


def estimate_pi(n_points, seed, num_procs):
    return run_chunks(split_work(n_points, seed, num_procs))


times = {}
for num_procs in (1, 4):    # USER: try different values to feel the difference
    start = time.perf_counter()
    pi_est = estimate_pi(N_POINTS, SEED, num_procs)
    times[num_procs] = time.perf_counter() - start
    print(f"pi ~ {pi_est:.5f}  ({times[num_procs]:.2f} sec with {num_procs} core(s))")

print(f"\nspeedup with 4 cores: {times[1] / times[4]:.1f}x (ideally: 4x)")
pi ~ 3.14162  (0.13 sec with 1 core(s))
pi ~ 3.14089  (0.13 sec with 4 core(s))

speedup with 4 cores: 1.0x (ideally: 4x)

Step 3: The Same Study in QUEENS#

Now we hand the orchestration to QUEENS.

  • num_samples - how many samples to run in total.

  • num_jobs - how many of those samples run at the same time.

  • num_procs - how many cores each job gets.


Example:

num_samples = 8 with num_jobs = 4 and num_procs = 1 means that four jobs run at the same time with one core each until eight samples are done.


Reading the timings: several outputs show up, each measuring something different:

  • Time for CALCULATION: … s (QUEENS) - wall-clock time of the calculation.

  • … it/s (QUEENS progress bar) - throughput: samples finished per second of that calculation time.

  • … s per sample (our own print) - the compute time of a single sample inside the driver.

  • … s for the whole study (our own print) - setup (start the worker pool) + calculation (the number QUEENS reports above) + teardown (save results, stop the workers).

[4]:
class PiDriver(Driver):
    """One QUEENS job = one pi estimate, spending its `num_procs` on the workers."""

    def __init__(self, parameters, n_points):
        super().__init__(parameters=parameters)
        self.n_points = n_points

    def run(self, sample, job_id, num_procs, experiment_dir, experiment_name):
        # the sampled float in [0, 1] becomes this job's integer seed
        seed = int(self.parameters.sample_as_dict(sample)["seed"] * 1e9)
        start = time.perf_counter()
        pi_est = estimate_pi(self.n_points, seed, num_procs)
        elapsed = time.perf_counter() - start
        return {"result": np.array([pi_est]), "time": np.array([elapsed])}


parameters = Parameters(seed=Uniform(lower_bound=0.0, upper_bound=1.0))  # each job's only input


def run_study(experiment_name, driver, num_jobs, num_procs, num_samples):
    """Run one MonteCarlo study under a given (num_jobs, num_procs) allocation."""
    with GlobalSettings(experiment_name=experiment_name, output_dir="./output") as gs:
        scheduler = Local(
            experiment_name=gs.experiment_name,
            num_jobs=num_jobs,
            num_procs=num_procs,
            overwrite_existing_experiment=True,
        )
        model = Simulation(scheduler=scheduler, driver=driver)
        iterator = MonteCarlo(
            model=model,
            parameters=parameters,
            global_settings=gs,
            seed=SEED,
            num_samples=num_samples,
            result_description={"write_results": True, "plot_results": False},
        )
        run_iterator(iterator, global_settings=gs)
        return load_result(gs.result_file(".pickle"))["raw_output_data"]


def report(num_jobs, num_procs, num_samples):
    """Run one allocation and print per-sample and whole-study timings."""
    start = time.perf_counter()
    out = run_study(
        experiment_name=f"PI_EXPERIMENT_{num_jobs}_jobs_{num_procs}_procs",
        driver=PiDriver(parameters, n_points=N_POINTS),
        num_jobs=num_jobs,
        num_procs=num_procs,
        num_samples=num_samples,
    )
    study_time = time.perf_counter() - start
    estimates = np.array(out["result"]).ravel()
    times = np.array(out["time"]).ravel()
    for i, (pi_i, t_i) in enumerate(zip(estimates, times), start=1):
        print(f"Sample {i}: pi ~ {pi_i:.5f}  ({t_i:.2f} sec with {num_procs} core(s))")
    print(f"\nmean pi ~ {estimates.mean():.5f} over {len(estimates)} samples")
    print(f"mean time ~ {times.mean():.2f} sec per sample")
    print(f"{study_time:.2f} sec for the whole study (setup + calculation + teardown)\n")


for num_jobs, num_procs in [(4, 1), (1, 4)]:    # USER: try (2, 2), or raise num_samples; try different values to feel the difference
    report(num_jobs=num_jobs, num_procs=num_procs, num_samples=8)


                                                 .**.
                                                 I  I
                                                 *  *
                                                :.  .:
                                                I    I
                                 :::           .*    *.           :*:
                                 I  *          *.    .*          *  I
                                .:   *:*::::   I      I   ::::*:*   :.
                                ::   :I:    ::.*      *.::    :I:   ::
                                :.   * *:    .V.      .V.    :* *   .:
                                :.  I   ::    I*.     *I    ::   I  .*
                                *. ::    .*  :: ::  :* ::  *:    :* .*
                                *. I       * I   :**:   I *       I .*
                                *.:.        I*    **    *I        .:.*
                                *:I        :*.* ::  *: *.*:        I.*
                                *I:       :*  .**    *I.  *:       :**
                                *V       *. .*.  *II*  .*. .*       V*
                                ** ..:*I***I*::::    ::::*I***I*:.. **
                                 ......                        ......


     :*IV$$$V*:        VV:        *VV    VVVVVVVVVVVF   *VVVVVVVVVVV.  .VF.        :VI     :FV$$$V*:
   *$$*:.  .:*V$*      $$:        *$V    $$*.........   *$I.........   .$$$*       *$V    V$F.  .:FV.
  V$*          *$$.    $$:        *$V    $$:            *$F            .$$F$V.     *$V   .$$.
 V$F            *$V    $$:        *$V    $$:            *$I            .$$ .V$*    *$V    F$$*:.
 $$:            :$$    $$:        *$V    $$$VVVVVVVV    *$$VVVVVVVV:   .$$   *$V.  *$V     .*FV$$V*.
 I$F        **  *$V    $$:        *$V    $$:            *$F            .$$    .I$* *$V          .*$$*
  V$*       :V$F$$.    I$F        V$*    $$:            *$F            .$$      :$$I$V            *$$
   *$$*:.  .:*$$$F      F$V*....*V$*     $$*.........   *$I.........   .$$        F$$V   V$*:   .:V$*
     :*IV$$VI*: :I:      .*FVVVVF:       VVVVVVVVVVVV   *VVVVVVVVVVV.  .VV         :VI    :*VV$$VI*.


                 QUEENS (Quantification of Uncertain Effects in ENgineering Systems):
                        a Python framework for solver-independent multi-query
                            analyses of large-scale computational models.


+---------------------------------------------------------------------------------------------------+
|                                               Local                                               |
|- - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - -|
| self                          : <queens.schedulers.local.Local object at 0x7f73422d00b0>          |
| experiment_name               : 'PI_EXPERIMENT_4_jobs_1_procs'                                    |
| num_jobs                      : 4                                                                 |
| num_procs                     : 1                                                                 |
| restart_workers               : False                                                             |
| verbose                       : True                                                              |
| experiment_base_dir           : None                                                              |
| overwrite_existing_experiment : True                                                              |
+---------------------------------------------------------------------------------------------------+

To view the Dask dashboard open this link in your browser: http://127.0.0.1:8787/status

+---------------------------------------------------------------------------------------------------+
|                                            Simulation                                             |
|- - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - -|
| self      : <queens.models.simulation.Simulation object at 0x7f7342108710>                        |
| scheduler : <queens.schedulers.local.Local object at 0x7f73422d00b0>                              |
| driver    : <__main__.PiDriver object at 0x7f7341bffbc0>                                          |
+---------------------------------------------------------------------------------------------------+


+---------------------------------------------------------------------------------------------------+
|                                            MonteCarlo                                             |
|- - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - -|
| self               : <queens.iterators.monte_carlo.MonteCarlo object at 0x7f7341b9e000>           |
| model              : <queens.models.simulation.Simulation object at 0x7f7342108710>               |
| parameters         : <queens.parameters.parameters.Parameters object at 0x7f7342b25c10>           |
| global_settings    : <queens.global_settings.GlobalSettings object at 0x7f7341bfefc0>             |
| seed               : 17                                                                           |
| num_samples        : 8                                                                            |
| result_description : {'write_results': True, 'plot_results': False}                               |
+---------------------------------------------------------------------------------------------------+


+---------------------------------------------------------------------------------------------------+
|                                          git information                                          |
|- - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - -|
| commit hash        : 78466880c472fd6c4f28673fda0589f0e934912f                                     |
| branch             : main                                                                         |
| clean working tree : False                                                                        |
+---------------------------------------------------------------------------------------------------+

MonteCarlo for experiment: PI_EXPERIMENT_4_jobs_1_procs

Starting Analysis...

 62%|██████▎   | 5/8 [00:01<00:00,  5.77it/s]

+---------------------------------------------------------------------------------------------------+
|                                   Batch summary for jobs 0 - 7                                    |
|- - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - -|
| number of jobs                : 8                                                                 |
| number of parallel jobs       : 4                                                                 |
| number of procs               : 1                                                                 |
| total elapsed time            : 1.078e+00s                                                        |
| average time per parallel job : 5.388e-01s                                                        |
+---------------------------------------------------------------------------------------------------+

100%|██████████| 8/8 [00:01<00:00,  7.42it/s]

Time for CALCULATION: 1.0830695629119873 s


Sample 1: pi ~ 3.14146  (0.33 sec with 1 core(s))
Sample 2: pi ~ 3.14130  (0.32 sec with 1 core(s))
Sample 3: pi ~ 3.14125  (0.32 sec with 1 core(s))
Sample 4: pi ~ 3.14138  (0.31 sec with 1 core(s))
Sample 5: pi ~ 3.14061  (0.22 sec with 1 core(s))
Sample 6: pi ~ 3.14154  (0.23 sec with 1 core(s))
Sample 7: pi ~ 3.14139  (0.23 sec with 1 core(s))
Sample 8: pi ~ 3.14197  (0.21 sec with 1 core(s))

mean pi ~ 3.14136 over 8 samples
mean time ~ 0.27 sec per sample
2.18 sec for the whole study (setup + calculation + teardown)



                                                 .**.
                                                 I  I
                                                 *  *
                                                :.  .:
                                                I    I
                                 :::           .*    *.           :*:
                                 I  *          *.    .*          *  I
                                .:   *:*::::   I      I   ::::*:*   :.
                                ::   :I:    ::.*      *.::    :I:   ::
                                :.   * *:    .V.      .V.    :* *   .:
                                :.  I   ::    I*.     *I    ::   I  .*
                                *. ::    .*  :: ::  :* ::  *:    :* .*
                                *. I       * I   :**:   I *       I .*
                                *.:.        I*    **    *I        .:.*
                                *:I        :*.* ::  *: *.*:        I.*
                                *I:       :*  .**    *I.  *:       :**
                                *V       *. .*.  *II*  .*. .*       V*
                                ** ..:*I***I*::::    ::::*I***I*:.. **
                                 ......                        ......


     :*IV$$$V*:        VV:        *VV    VVVVVVVVVVVF   *VVVVVVVVVVV.  .VF.        :VI     :FV$$$V*:
   *$$*:.  .:*V$*      $$:        *$V    $$*.........   *$I.........   .$$$*       *$V    V$F.  .:FV.
  V$*          *$$.    $$:        *$V    $$:            *$F            .$$F$V.     *$V   .$$.
 V$F            *$V    $$:        *$V    $$:            *$I            .$$ .V$*    *$V    F$$*:.
 $$:            :$$    $$:        *$V    $$$VVVVVVVV    *$$VVVVVVVV:   .$$   *$V.  *$V     .*FV$$V*.
 I$F        **  *$V    $$:        *$V    $$:            *$F            .$$    .I$* *$V          .*$$*
  V$*       :V$F$$.    I$F        V$*    $$:            *$F            .$$      :$$I$V            *$$
   *$$*:.  .:*$$$F      F$V*....*V$*     $$*.........   *$I.........   .$$        F$$V   V$*:   .:V$*
     :*IV$$VI*: :I:      .*FVVVVF:       VVVVVVVVVVVV   *VVVVVVVVVVV.  .VV         :VI    :*VV$$VI*.


                 QUEENS (Quantification of Uncertain Effects in ENgineering Systems):
                        a Python framework for solver-independent multi-query
                            analyses of large-scale computational models.


+---------------------------------------------------------------------------------------------------+
|                                               Local                                               |
|- - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - -|
| self                          : <queens.schedulers.local.Local object at 0x7f733bfe90a0>          |
| experiment_name               : 'PI_EXPERIMENT_1_jobs_4_procs'                                    |
| num_jobs                      : 1                                                                 |
| num_procs                     : 4                                                                 |
| restart_workers               : False                                                             |
| verbose                       : True                                                              |
| experiment_base_dir           : None                                                              |
| overwrite_existing_experiment : True                                                              |
+---------------------------------------------------------------------------------------------------+

To view the Dask dashboard open this link in your browser: http://127.0.0.1:8787/status

+---------------------------------------------------------------------------------------------------+
|                                            Simulation                                             |
|- - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - -|
| self      : <queens.models.simulation.Simulation object at 0x7f733bfe97f0>                        |
| scheduler : <queens.schedulers.local.Local object at 0x7f733bfe90a0>                              |
| driver    : <__main__.PiDriver object at 0x7f7341cde450>                                          |
+---------------------------------------------------------------------------------------------------+


+---------------------------------------------------------------------------------------------------+
|                                            MonteCarlo                                             |
|- - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - -|
| self               : <queens.iterators.monte_carlo.MonteCarlo object at 0x7f733bed8800>           |
| model              : <queens.models.simulation.Simulation object at 0x7f733bfe97f0>               |
| parameters         : <queens.parameters.parameters.Parameters object at 0x7f7342b25c10>           |
| global_settings    : <queens.global_settings.GlobalSettings object at 0x7f73422d00b0>             |
| seed               : 17                                                                           |
| num_samples        : 8                                                                            |
| result_description : {'write_results': True, 'plot_results': False}                               |
+---------------------------------------------------------------------------------------------------+


+---------------------------------------------------------------------------------------------------+
|                                          git information                                          |
|- - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - -|
| commit hash        : 78466880c472fd6c4f28673fda0589f0e934912f                                     |
| branch             : main                                                                         |
| clean working tree : False                                                                        |
+---------------------------------------------------------------------------------------------------+

MonteCarlo for experiment: PI_EXPERIMENT_1_jobs_4_procs

Starting Analysis...

100%|██████████| 8/8 [00:01<00:00,  7.32it/s]

+---------------------------------------------------------------------------------------------------+
|                                   Batch summary for jobs 0 - 7                                    |
|- - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - -|
| number of jobs                : 8                                                                 |
| number of parallel jobs       : 1                                                                 |
| number of procs               : 4                                                                 |
| total elapsed time            : 1.377e+00s                                                        |
| average time per parallel job : 1.722e-01s                                                        |
+---------------------------------------------------------------------------------------------------+

100%|██████████| 8/8 [00:01<00:00,  5.81it/s]

Time for CALCULATION: 1.3819146156311035 s


Sample 1: pi ~ 3.14131  (0.13 sec with 4 core(s))
Sample 2: pi ~ 3.14146  (0.13 sec with 4 core(s))
Sample 3: pi ~ 3.14214  (0.13 sec with 4 core(s))
Sample 4: pi ~ 3.14210  (0.13 sec with 4 core(s))
Sample 5: pi ~ 3.14172  (0.13 sec with 4 core(s))
Sample 6: pi ~ 3.14094  (0.13 sec with 4 core(s))
Sample 7: pi ~ 3.14090  (0.13 sec with 4 core(s))
Sample 8: pi ~ 3.14103  (0.13 sec with 4 core(s))

mean pi ~ 3.14145 over 8 samples
mean time ~ 0.13 sec per sample
1.93 sec for the whole study (setup + calculation + teardown)

Step 4: Jobs with Phases#

Real jobs are rarely pure heavy computation - they run in phases. A simulation first does cheap serial preprocessing (meshing, reading input files) on one core, then the expensive solver runs on many, and at the end, a light postprocessing step is serial again. A QUEENS job holds one allocation from start to finish, so num_procs must be sized for the hungriest phase - the solver - and during every cheap phase, most of those cores sit idle.

TwoPhaseDriver mimics the first two phases with the two parts of estimate_pi():

  • preparation (e.g., meshing): split_work(), stretched by a PREPARATION_SECONDS sleep to give it a realistic weight - serial, uses one core.

  • computation (e.g., simulation): run_chunks() on num_procs cores.

Postprocessing would simply be a third phase - everything below extends to any number of phases.

[5]:
PREPARATION_SECONDS = 1.0   # cost of preparation
SAMPLES = 8                 # USER: try different values to feel the difference
PROCS = 4                   # USER: try different values to feel the difference
JOBS = 1                    # USER: try different values to feel the difference


class TwoPhaseDriver(Driver):
    """One job = serial preparation (1 core) + parallel computation (`num_procs` cores)."""

    def __init__(self, parameters, n_points):
        super().__init__(parameters=parameters)
        self.n_points = n_points

    def run(self, sample, job_id, num_procs, experiment_dir, experiment_name):
        seed = int(self.parameters.sample_as_dict(sample)["seed"] * 1e9)
        t0 = time.perf_counter()
        time.sleep(PREPARATION_SECONDS)     # serial preparation (1 core)
        chunks = split_work(self.n_points, seed, num_procs)
        t1 = time.perf_counter()
        pi = run_chunks(chunks)             # parallel computation (PROCS cores)
        t2 = time.perf_counter()
        return {
            "result": np.array([pi]),
            "preparation": np.array([t1 - t0]),
            "computation": np.array([t2 - t1]),
        }


start = time.perf_counter()
out = run_study(
    experiment_name="PI_EXPERIMENT_1_driver",
    driver=TwoPhaseDriver(parameters, N_POINTS),
    num_jobs=JOBS,
    num_procs=PROCS,
    num_samples=SAMPLES,
)
one_driver_time = time.perf_counter() - start  # compared against the two-driver version below

estimates = np.array(out["result"]).ravel()
preparation = np.array(out["preparation"]).ravel()
computation = np.array(out["computation"]).ravel()
times = preparation + computation

for i, (pi_i, p_i, c_i) in enumerate(zip(estimates, preparation, computation), start=1):
    print(f"Sample {i}: pi ~ {pi_i:.5f}  ({p_i:.2f} sec preparation + {c_i:.2f} sec computation with {PROCS} core(s))")

idle = (PROCS - 1) * preparation.sum()  # core-seconds idling during preparation
allocated = PROCS * times.sum()         # core-seconds the allocation holds in total

print(f"\nmean pi ~ {estimates.mean():.5f} over {len(estimates)} samples")
print(f"mean time ~ {preparation.mean():.2f} sec preparation + {computation.mean():.2f} sec computation per sample\n")
print(f"{one_driver_time:.2f} sec for the whole study (setup + calculation + teardown)\n")
print(f"Total allocated time ~ {allocated:.2f} sec (total compute time)")
print(f"Total idle time ~ {idle:.2f} sec")
print(f"Unused compute time (idle/allocated) ~ {idle / allocated:.0%}")


                                                 .**.
                                                 I  I
                                                 *  *
                                                :.  .:
                                                I    I
                                 :::           .*    *.           :*:
                                 I  *          *.    .*          *  I
                                .:   *:*::::   I      I   ::::*:*   :.
                                ::   :I:    ::.*      *.::    :I:   ::
                                :.   * *:    .V.      .V.    :* *   .:
                                :.  I   ::    I*.     *I    ::   I  .*
                                *. ::    .*  :: ::  :* ::  *:    :* .*
                                *. I       * I   :**:   I *       I .*
                                *.:.        I*    **    *I        .:.*
                                *:I        :*.* ::  *: *.*:        I.*
                                *I:       :*  .**    *I.  *:       :**
                                *V       *. .*.  *II*  .*. .*       V*
                                ** ..:*I***I*::::    ::::*I***I*:.. **
                                 ......                        ......


     :*IV$$$V*:        VV:        *VV    VVVVVVVVVVVF   *VVVVVVVVVVV.  .VF.        :VI     :FV$$$V*:
   *$$*:.  .:*V$*      $$:        *$V    $$*.........   *$I.........   .$$$*       *$V    V$F.  .:FV.
  V$*          *$$.    $$:        *$V    $$:            *$F            .$$F$V.     *$V   .$$.
 V$F            *$V    $$:        *$V    $$:            *$I            .$$ .V$*    *$V    F$$*:.
 $$:            :$$    $$:        *$V    $$$VVVVVVVV    *$$VVVVVVVV:   .$$   *$V.  *$V     .*FV$$V*.
 I$F        **  *$V    $$:        *$V    $$:            *$F            .$$    .I$* *$V          .*$$*
  V$*       :V$F$$.    I$F        V$*    $$:            *$F            .$$      :$$I$V            *$$
   *$$*:.  .:*$$$F      F$V*....*V$*     $$*.........   *$I.........   .$$        F$$V   V$*:   .:V$*
     :*IV$$VI*: :I:      .*FVVVVF:       VVVVVVVVVVVV   *VVVVVVVVVVV.  .VV         :VI    :*VV$$VI*.


                 QUEENS (Quantification of Uncertain Effects in ENgineering Systems):
                        a Python framework for solver-independent multi-query
                            analyses of large-scale computational models.


+---------------------------------------------------------------------------------------------------+
|                                               Local                                               |
|- - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - -|
| self                          : <queens.schedulers.local.Local object at 0x7f73422d00b0>          |
| experiment_name               : 'PI_EXPERIMENT_1_driver'                                          |
| num_jobs                      : 1                                                                 |
| num_procs                     : 4                                                                 |
| restart_workers               : False                                                             |
| verbose                       : True                                                              |
| experiment_base_dir           : None                                                              |
| overwrite_existing_experiment : True                                                              |
+---------------------------------------------------------------------------------------------------+

To view the Dask dashboard open this link in your browser: http://127.0.0.1:8787/status

+---------------------------------------------------------------------------------------------------+
|                                            Simulation                                             |
|- - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - -|
| self      : <queens.models.simulation.Simulation object at 0x7f7341cde450>                        |
| scheduler : <queens.schedulers.local.Local object at 0x7f73422d00b0>                              |
| driver    : <__main__.TwoPhaseDriver object at 0x7f7384fdf6b0>                                    |
+---------------------------------------------------------------------------------------------------+


+---------------------------------------------------------------------------------------------------+
|                                            MonteCarlo                                             |
|- - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - -|
| self               : <queens.iterators.monte_carlo.MonteCarlo object at 0x7f733be6f3b0>           |
| model              : <queens.models.simulation.Simulation object at 0x7f7341cde450>               |
| parameters         : <queens.parameters.parameters.Parameters object at 0x7f7342b25c10>           |
| global_settings    : <queens.global_settings.GlobalSettings object at 0x7f7384fde7b0>             |
| seed               : 17                                                                           |
| num_samples        : 8                                                                            |
| result_description : {'write_results': True, 'plot_results': False}                               |
+---------------------------------------------------------------------------------------------------+


+---------------------------------------------------------------------------------------------------+
|                                          git information                                          |
|- - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - -|
| commit hash        : 78466880c472fd6c4f28673fda0589f0e934912f                                     |
| branch             : main                                                                         |
| clean working tree : False                                                                        |
+---------------------------------------------------------------------------------------------------+

MonteCarlo for experiment: PI_EXPERIMENT_1_driver

Starting Analysis...

100%|██████████| 8/8 [00:09<00:00,  1.14s/it]

+---------------------------------------------------------------------------------------------------+
|                                   Batch summary for jobs 0 - 7                                    |
|- - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - -|
| number of jobs                : 8                                                                 |
| number of parallel jobs       : 1                                                                 |
| number of procs               : 4                                                                 |
| total elapsed time            : 9.368e+00s                                                        |
| average time per parallel job : 1.171e+00s                                                        |
+---------------------------------------------------------------------------------------------------+

100%|██████████| 8/8 [00:09<00:00,  1.17s/it]

Time for CALCULATION: 9.373283624649048 s


Sample 1: pi ~ 3.14131  (1.00 sec preparation + 0.13 sec computation with 4 core(s))
Sample 2: pi ~ 3.14146  (1.00 sec preparation + 0.12 sec computation with 4 core(s))
Sample 3: pi ~ 3.14214  (1.00 sec preparation + 0.13 sec computation with 4 core(s))
Sample 4: pi ~ 3.14210  (1.00 sec preparation + 0.13 sec computation with 4 core(s))
Sample 5: pi ~ 3.14172  (1.00 sec preparation + 0.14 sec computation with 4 core(s))
Sample 6: pi ~ 3.14094  (1.00 sec preparation + 0.13 sec computation with 4 core(s))
Sample 7: pi ~ 3.14090  (1.00 sec preparation + 0.13 sec computation with 4 core(s))
Sample 8: pi ~ 3.14103  (1.00 sec preparation + 0.13 sec computation with 4 core(s))

mean pi ~ 3.14145 over 8 samples
mean time ~ 1.00 sec preparation + 0.13 sec computation per sample

9.89 sec for the whole study (setup + calculation + teardown)

Total allocated time ~ 36.11 sec (total compute time)
Total idle time ~ 24.00 sec
Unused compute time (idle/allocated) ~ 66%

The Fix: One Driver per Phase#

Nothing forces both phases to share one allocation. Split the driver into two, and give each its own study with its own allocation:

  1. PreparationDriver builds the chunk plan (split_work()) and writes it to the disk.

  2. ComputationDriver reads the plan and runs it (run_chunks()).

Both studies use the same iterator seed, so the computation study sees the same samples - and the same job_ids - as the preparation study; that is how each computation job finds its plan file.

[6]:
PLAN_DIR = Path("./output/pi_plans")  # PreparationDriver writes plans here, ComputationDriver reads them

PREPARATION_SECONDS = 1.0   # cost of preparation
PREPARATION_JOBS = 4        # USER: try different values to feel the difference
PREPARATION_PROCS = 1       # USER: try different values to feel the difference

COMPUTATION_JOBS = 1        # USER: try different values to feel the difference
COMPUTATION_PROCS = 4       # USER: try different values to feel the difference

SAMPLES = 8                 # USER: try different values to feel the difference


class PreparationDriver(Driver):
    """Serial phase only: build the chunk plan, write it to disk."""

    def __init__(self, parameters, n_points):
        super().__init__(parameters=parameters)
        self.n_points = n_points

    def run(self, sample, job_id, num_procs, experiment_dir, experiment_name):
        seed = int(self.parameters.sample_as_dict(sample)["seed"] * 1e9)
        start = time.perf_counter()
        time.sleep(PREPARATION_SECONDS)
        chunks = split_work(self.n_points, seed, COMPUTATION_PROCS)  # plan for the computation study
        np.save(PLAN_DIR / f"plan_{job_id}.npy", np.array(chunks))
        elapsed = time.perf_counter() - start
        return {"result": np.array([float(len(chunks))]), "time": np.array([elapsed])}


class ComputationDriver(Driver):
    """Parallel phase only: read the plan back, run the chunks on `num_procs` cores."""

    def run(self, sample, job_id, num_procs, experiment_dir, experiment_name):
        start = time.perf_counter()
        chunks = np.load(PLAN_DIR / f"plan_{job_id}.npy")
        pi = run_chunks([(int(n), int(s)) for n, s in chunks])
        elapsed = time.perf_counter() - start
        return {"result": np.array([pi]), "time": np.array([elapsed])}


PLAN_DIR.mkdir(parents=True, exist_ok=True)

start = time.perf_counter()
out_preparation = run_study(
    experiment_name="PI_EXPERIMENT_PREPARATION",
    driver=PreparationDriver(parameters, N_POINTS),
    num_jobs=PREPARATION_JOBS,
    num_procs=PREPARATION_PROCS,
    num_samples=SAMPLES,
)
preparation_time = time.perf_counter() - start

start = time.perf_counter()
out_computation = run_study(
    experiment_name="PI_EXPERIMENT_COMPUTATION",
    driver=ComputationDriver(parameters),
    num_jobs=COMPUTATION_JOBS,
    num_procs=COMPUTATION_PROCS,
    num_samples=SAMPLES,
)
computation_time = time.perf_counter() - start
two_driver_time = preparation_time + computation_time

print(f"Preparation study ({PREPARATION_JOBS} job(s) x {PREPARATION_PROCS} core(s)):")
print("-" * 50)
chunk_counts = np.array(out_preparation["result"]).ravel()
times = np.array(out_preparation["time"]).ravel()
# for i, (n_i, t_i) in enumerate(zip(chunk_counts, times), start=1): # USER: uncomment if needed
#     print(f"Sample {i}: Split into {int(n_i)} chunks  ({t_i:.2f} sec with {PREPARATION_PROCS} core(s))")
print(f"mean time ~ {times.mean():.2f} sec per sample")
print(f"{preparation_time:.2f} sec for the whole study (setup + calculation + teardown)\n")

print(f"Computation study ({COMPUTATION_JOBS} job(s) x {COMPUTATION_PROCS} core(s)):")
print("-" * 50)
estimates = np.array(out_computation["result"]).ravel()
times = np.array(out_computation["time"]).ravel()
for i, (pi_i, t_i) in enumerate(zip(estimates, times), start=1):
    print(f"Sample {i}: pi ~ {pi_i:.5f}  ({t_i:.2f} sec with {COMPUTATION_PROCS} core(s))")
print(f"\nmean pi ~ {estimates.mean():.5f} over {len(estimates)} samples")
print(f"mean time ~ {times.mean():.2f} sec per sample\n")
print(f"{computation_time:.2f} sec for the whole study (setup + calculation + teardown)\n")

print(f"One driver ('TwoPhaseDriver') ~ {one_driver_time:.2f} sec (see the previous output)")
print(f"Two drivers ('PreparationDriver' = {preparation_time:.2f} sec + 'ComputationDriver' = {computation_time:.2f} sec) ~ {two_driver_time:.2f} sec")
print(f"Splitting up the study into multiple drivers is {one_driver_time / two_driver_time:.1f}x faster!")


                                                 .**.
                                                 I  I
                                                 *  *
                                                :.  .:
                                                I    I
                                 :::           .*    *.           :*:
                                 I  *          *.    .*          *  I
                                .:   *:*::::   I      I   ::::*:*   :.
                                ::   :I:    ::.*      *.::    :I:   ::
                                :.   * *:    .V.      .V.    :* *   .:
                                :.  I   ::    I*.     *I    ::   I  .*
                                *. ::    .*  :: ::  :* ::  *:    :* .*
                                *. I       * I   :**:   I *       I .*
                                *.:.        I*    **    *I        .:.*
                                *:I        :*.* ::  *: *.*:        I.*
                                *I:       :*  .**    *I.  *:       :**
                                *V       *. .*.  *II*  .*. .*       V*
                                ** ..:*I***I*::::    ::::*I***I*:.. **
                                 ......                        ......


     :*IV$$$V*:        VV:        *VV    VVVVVVVVVVVF   *VVVVVVVVVVV.  .VF.        :VI     :FV$$$V*:
   *$$*:.  .:*V$*      $$:        *$V    $$*.........   *$I.........   .$$$*       *$V    V$F.  .:FV.
  V$*          *$$.    $$:        *$V    $$:            *$F            .$$F$V.     *$V   .$$.
 V$F            *$V    $$:        *$V    $$:            *$I            .$$ .V$*    *$V    F$$*:.
 $$:            :$$    $$:        *$V    $$$VVVVVVVV    *$$VVVVVVVV:   .$$   *$V.  *$V     .*FV$$V*.
 I$F        **  *$V    $$:        *$V    $$:            *$F            .$$    .I$* *$V          .*$$*
  V$*       :V$F$$.    I$F        V$*    $$:            *$F            .$$      :$$I$V            *$$
   *$$*:.  .:*$$$F      F$V*....*V$*     $$*.........   *$I.........   .$$        F$$V   V$*:   .:V$*
     :*IV$$VI*: :I:      .*FVVVVF:       VVVVVVVVVVVV   *VVVVVVVVVVV.  .VV         :VI    :*VV$$VI*.


                 QUEENS (Quantification of Uncertain Effects in ENgineering Systems):
                        a Python framework for solver-independent multi-query
                            analyses of large-scale computational models.


+---------------------------------------------------------------------------------------------------+
|                                               Local                                               |
|- - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - -|
| self                          : <queens.schedulers.local.Local object at 0x7f73422d00b0>          |
| experiment_name               : 'PI_EXPERIMENT_PREPARATION'                                       |
| num_jobs                      : 4                                                                 |
| num_procs                     : 1                                                                 |
| restart_workers               : False                                                             |
| verbose                       : True                                                              |
| experiment_base_dir           : None                                                              |
| overwrite_existing_experiment : True                                                              |
+---------------------------------------------------------------------------------------------------+

To view the Dask dashboard open this link in your browser: http://127.0.0.1:8787/status

+---------------------------------------------------------------------------------------------------+
|                                            Simulation                                             |
|- - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - -|
| self      : <queens.models.simulation.Simulation object at 0x7f7341bffa40>                        |
| scheduler : <queens.schedulers.local.Local object at 0x7f73422d00b0>                              |
| driver    : <__main__.PreparationDriver object at 0x7f7341bffc20>                                 |
+---------------------------------------------------------------------------------------------------+


+---------------------------------------------------------------------------------------------------+
|                                            MonteCarlo                                             |
|- - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - -|
| self               : <queens.iterators.monte_carlo.MonteCarlo object at 0x7f733bdafda0>           |
| model              : <queens.models.simulation.Simulation object at 0x7f7341bffa40>               |
| parameters         : <queens.parameters.parameters.Parameters object at 0x7f7342b25c10>           |
| global_settings    : <queens.global_settings.GlobalSettings object at 0x7f7341bfe4b0>             |
| seed               : 17                                                                           |
| num_samples        : 8                                                                            |
| result_description : {'write_results': True, 'plot_results': False}                               |
+---------------------------------------------------------------------------------------------------+


+---------------------------------------------------------------------------------------------------+
|                                          git information                                          |
|- - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - -|
| commit hash        : 78466880c472fd6c4f28673fda0589f0e934912f                                     |
| branch             : main                                                                         |
| clean working tree : False                                                                        |
+---------------------------------------------------------------------------------------------------+

MonteCarlo for experiment: PI_EXPERIMENT_PREPARATION

Starting Analysis...

 62%|██████▎   | 5/8 [00:02<00:01,  2.29it/s]

+---------------------------------------------------------------------------------------------------+
|                                   Batch summary for jobs 0 - 7                                    |
|- - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - -|
| number of jobs                : 8                                                                 |
| number of parallel jobs       : 4                                                                 |
| number of procs               : 1                                                                 |
| total elapsed time            : 2.521e+00s                                                        |
| average time per parallel job : 1.261e+00s                                                        |
+---------------------------------------------------------------------------------------------------+

100%|██████████| 8/8 [00:02<00:00,  3.17it/s]

Time for CALCULATION: 2.5261387825012207 s




                                                 .**.
                                                 I  I
                                                 *  *
                                                :.  .:
                                                I    I
                                 :::           .*    *.           :*:
                                 I  *          *.    .*          *  I
                                .:   *:*::::   I      I   ::::*:*   :.
                                ::   :I:    ::.*      *.::    :I:   ::
                                :.   * *:    .V.      .V.    :* *   .:
                                :.  I   ::    I*.     *I    ::   I  .*
                                *. ::    .*  :: ::  :* ::  *:    :* .*
                                *. I       * I   :**:   I *       I .*
                                *.:.        I*    **    *I        .:.*
                                *:I        :*.* ::  *: *.*:        I.*
                                *I:       :*  .**    *I.  *:       :**
                                *V       *. .*.  *II*  .*. .*       V*
                                ** ..:*I***I*::::    ::::*I***I*:.. **
                                 ......                        ......


     :*IV$$$V*:        VV:        *VV    VVVVVVVVVVVF   *VVVVVVVVVVV.  .VF.        :VI     :FV$$$V*:
   *$$*:.  .:*V$*      $$:        *$V    $$*.........   *$I.........   .$$$*       *$V    V$F.  .:FV.
  V$*          *$$.    $$:        *$V    $$:            *$F            .$$F$V.     *$V   .$$.
 V$F            *$V    $$:        *$V    $$:            *$I            .$$ .V$*    *$V    F$$*:.
 $$:            :$$    $$:        *$V    $$$VVVVVVVV    *$$VVVVVVVV:   .$$   *$V.  *$V     .*FV$$V*.
 I$F        **  *$V    $$:        *$V    $$:            *$F            .$$    .I$* *$V          .*$$*
  V$*       :V$F$$.    I$F        V$*    $$:            *$F            .$$      :$$I$V            *$$
   *$$*:.  .:*$$$F      F$V*....*V$*     $$*.........   *$I.........   .$$        F$$V   V$*:   .:V$*
     :*IV$$VI*: :I:      .*FVVVVF:       VVVVVVVVVVVV   *VVVVVVVVVVV.  .VV         :VI    :*VV$$VI*.


                 QUEENS (Quantification of Uncertain Effects in ENgineering Systems):
                        a Python framework for solver-independent multi-query
                            analyses of large-scale computational models.


+---------------------------------------------------------------------------------------------------+
|                                               Local                                               |
|- - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - -|
| self                          : <queens.schedulers.local.Local object at 0x7f733bd62ed0>          |
| experiment_name               : 'PI_EXPERIMENT_COMPUTATION'                                       |
| num_jobs                      : 1                                                                 |
| num_procs                     : 4                                                                 |
| restart_workers               : False                                                             |
| verbose                       : True                                                              |
| experiment_base_dir           : None                                                              |
| overwrite_existing_experiment : True                                                              |
+---------------------------------------------------------------------------------------------------+

To view the Dask dashboard open this link in your browser: http://127.0.0.1:8787/status

+---------------------------------------------------------------------------------------------------+
|                                            Simulation                                             |
|- - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - -|
| self      : <queens.models.simulation.Simulation object at 0x7f733bd601a0>                        |
| scheduler : <queens.schedulers.local.Local object at 0x7f733bd62ed0>                              |
| driver    : <__main__.ComputationDriver object at 0x7f7341cde450>                                 |
+---------------------------------------------------------------------------------------------------+


+---------------------------------------------------------------------------------------------------+
|                                            MonteCarlo                                             |
|- - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - -|
| self               : <queens.iterators.monte_carlo.MonteCarlo object at 0x7f733bdad040>           |
| model              : <queens.models.simulation.Simulation object at 0x7f733bd601a0>               |
| parameters         : <queens.parameters.parameters.Parameters object at 0x7f7342b25c10>           |
| global_settings    : <queens.global_settings.GlobalSettings object at 0x7f73422d00b0>             |
| seed               : 17                                                                           |
| num_samples        : 8                                                                            |
| result_description : {'write_results': True, 'plot_results': False}                               |
+---------------------------------------------------------------------------------------------------+


+---------------------------------------------------------------------------------------------------+
|                                          git information                                          |
|- - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - -|
| commit hash        : 78466880c472fd6c4f28673fda0589f0e934912f                                     |
| branch             : main                                                                         |
| clean working tree : False                                                                        |
+---------------------------------------------------------------------------------------------------+

MonteCarlo for experiment: PI_EXPERIMENT_COMPUTATION

Starting Analysis...

100%|██████████| 8/8 [00:01<00:00,  7.24it/s]

+---------------------------------------------------------------------------------------------------+
|                                   Batch summary for jobs 0 - 7                                    |
|- - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - -|
| number of jobs                : 8                                                                 |
| number of parallel jobs       : 1                                                                 |
| number of procs               : 4                                                                 |
| total elapsed time            : 1.367e+00s                                                        |
| average time per parallel job : 1.709e-01s                                                        |
+---------------------------------------------------------------------------------------------------+

100%|██████████| 8/8 [00:01<00:00,  5.85it/s]

Time for CALCULATION: 1.3720757961273193 s


Preparation study (4 job(s) x 1 core(s)):
--------------------------------------------------
mean time ~ 1.00 sec per sample
3.31 sec for the whole study (setup + calculation + teardown)

Computation study (1 job(s) x 4 core(s)):
--------------------------------------------------
Sample 1: pi ~ 3.14131  (0.13 sec with 4 core(s))
Sample 2: pi ~ 3.14146  (0.13 sec with 4 core(s))
Sample 3: pi ~ 3.14214  (0.13 sec with 4 core(s))
Sample 4: pi ~ 3.14210  (0.13 sec with 4 core(s))
Sample 5: pi ~ 3.14172  (0.13 sec with 4 core(s))
Sample 6: pi ~ 3.14094  (0.13 sec with 4 core(s))
Sample 7: pi ~ 3.14090  (0.13 sec with 4 core(s))
Sample 8: pi ~ 3.14103  (0.13 sec with 4 core(s))

mean pi ~ 3.14145 over 8 samples
mean time ~ 0.13 sec per sample

1.86 sec for the whole study (setup + calculation + teardown)

One driver ('TwoPhaseDriver') ~ 9.89 sec (see the previous output)
Two drivers ('PreparationDriver' = 3.31 sec + 'ComputationDriver' = 1.86 sec) ~ 5.17 sec
Splitting up the study into multiple drivers is 1.9x faster!

Lessons Learned#

  • The computing power in use is the product of the two scheduler parameters: cores in use = num_jobs × num_procs. The product is bounded by the available cores of the computer; the split between the two factors is up to the user.

  • An allocation of computing power is fixed from start to finish. If the job runs in phases, num_procs must be sized for the most computationally demanding phase - and during every cheaper phase, the extra cores sit idle.

  • Splitting the study into one driver per phase gives control to split the computing power optimally - the idle time turns into throughput (compare the one-driver and two-driver timings above).

The same logic applies on a cluster, where num_nodes joins the split as a third factor (see The Parameters above).