A technical dice audit: 14.4 million faces and complete replay
A conditional proof of rejection sampling, 44 statistical tests, five charts, and every seed, raw outcome and script needed to audit Dicecore 3.7.1.
We collected 14,400,000 faces through the public @erpg/dicecore@3.7.1 API, retained all 428,000 seeds and replayed every outcome exactly. None of the 44 primary tests rejected its null model after Holm adjustment at a 1% family significance level. The installed engine was also checked against the official npm artifact: all 89 engine and license files match.
This is an internal engineering study with public data and code for external verification. Its conclusion is specific: the observations are compatible with uniform face probabilities and independence of the examined pairs. The integer-to-face mapping has a mathematical proof of absence of modulo bias conditional on uniform independent source words. A finite sample does not prove perfect randomness, and deterministic replay does not certify resistance to cheating.
Download the complete study, including raw faces and seeds, inspect the unrounded statistics or use the CSV results. The reproduction guide includes complete commands. All figures are available as SVG and PNG in the archive.
1. The execution path being audited
The frontend calls the public @erpg/dicecore/systems/mixed module with these options:
import { rollMixedDice } from "@erpg/dicecore/systems/mixed";
const result = rollMixedDice("1d20", {
detail: "compact",
randomAlgorithm: "xoshiro128ss",
});
console.log(result.dice.map((die) => die.rawValue));
The configuration name xoshiro128ss means *xoshiro128\\ 1.1, by Blackman and Vigna. It is explicitly selected by the site's rollCompactMixedDice wrapper. The library default when the option is absent is MT19937*. Both are measured separately; the library default should not be confused with the application's choice. Inspect the frontend wrapper at the examined revision and the source implementations of xoshiro, MT19937 and the execution context.
Without a supplied seed, the engine requests four 32-bit words from crypto.getRandomValues: 128 bits of seed material per call. Executions have isolated PRNG state. A 500d20 call draws its 500 faces from that state's sequence; a subsequent call starts with a fresh seed. If the cryptographic source is unavailable or fails, the implementation throws RNG_UNAVAILABLE rather than silently switching to Math.random. See replay.ts and the Web Crypto specification.
Cryptographic seed acquisition is different from a cryptographic generator. Xoshiro and MT19937 are deterministic, non-cryptographic PRNGs. Neither guarantees unpredictability to an adversary who knows or recovers its state. Generator state size and period do not create more entropy than the 128-bit seed supplies. Consult the authors' PRNG reference and Mersenne Twister project.
3D dice receive an already-computed result
The character-sheet path computes the roll before passing faces to the 3D adapter. rawValue is the sampled face; textures, materials and motion are presentation. We measure the API's raw faces, not mesh orientation or animation pixels. The relationship can be inspected in the sheet runtime and 3D adapter.
The backend/Fortuna has a Go implementation. This study does not measure that runtime, production-player histories, or every supported system and modifier. It examines the identified JavaScript package through the frontend's API and algorithm option, in Node and an actual browser.
2. Proof: why rejection mapping eliminates modulo bias
Let U be uniform on the integers {0, …, Q−1}, where Q=2^32. For an s-sided die, directly returning 1+(U mod s) is uniform only if s divides Q. Otherwise some residues have an additional preimage.
Both generators use this mapping, shown as pseudocode:
Q = 2^32
L = floor(Q / s) * s
repeat:
U = next UInt32 word
until U < L
face = 1 + (U mod s)
Write L=m·s. For face k, the accepted interval contains exactly m preimages: k−1, k−1+s, …, k−1+(m−1)s. Therefore:
P(face = k | U < L) = m/L = m/(m*s) = 1/s
Under uniform independent attempts, repeating until acceptance preserves that distribution. Rejection depends on the raw word, before a completed face exists. It is not filtering completed rolls to produce favorable outcomes or an attractive histogram.
For d20, 2^32 = 214,748,364 × 20 + 16, so:
L = 4,294,967,280
P(rejecting a word) = 16 / 4,294,967,296 ≈ 3.725290298 × 10^-9
E[attempts per face] = Q/L ≈ 1.000000003725
The mapping removes modulo bias by construction. The quality of the source words is a separate question: a deterministic finite-state PRNG is not an infinite stream of independent random variables.
Figure 1 — an exact teaching example, not collected observations. Eight bits provide 256 words. Direct modulo assigns faces 1–16 probability 13/256=5.078125%, and faces 17–20 12/256=4.6875%. Accepting only 0–239 gives every face 12/240=5%. The vertical axis starts at 4.4% to expose the difference; the real engine uses 32-bit words. Open the SVG.
3. A protocol fixed before analysis
The original protocol was frozen at 2026-09-28T19:54:04Z. Collection started at 19:56:32Z; merging the browser dataset finished at 20:01:48Z. These are recorded internal timestamps, not an externally notarized preregistration.
| Profile | Algorithm | Faces per die | Datasets | | --- | --- | ---: | ---: | | Node, calls of 500 dice | xoshiro128ss | 1,000,000 | 7 | | Node, calls of 500 dice | mt19937 | 1,000,000 | 7 | | Node, fresh seed per face | xoshiro128ss | 50,000 | 7 | | Chromium, fresh seed per face | xoshiro128ss, d20 | 50,000 | 1 |
The seven types are d4, d6, d8, d10, d12, d20, d100: 22 datasets and 14.4 million faces. Batch datasets contain 2,000 calls each. Formulas have no rerolls, explosions, keep/drop, bonuses or comparisons. Every rawValue is retained in execution order. No supplied seed was selected to produce favorable results.
The selection rule was to publish the first completed collection, including rejections. Small development fixtures are marked smokeTest and refused by the publication analysis. The collection was not repeated after examining statistical results. The protocol SHA-256 is:
b84de0bc80d31fe22cca382f74476551b927cebbc17f1d6b5e4992264e293843
The manifest records filenames, hashes, formulas, replay versions and environments. We used Node 24.18.0, Chromium 154.0.8037.0 in a secure localhost context, Python 3.12.14, NumPy 2.3.3, SciPy 1.16.3 and Matplotlib 3.10.7. The browser sample actually ran the JavaScript library with browser Web Crypto; it was not a Node simulation of a browser.
4. Tests and hypotheses
Face uniformity
For s sides, the null hypothesis is P(X=k)=1/s for every face. For N observations, observed counts O_k and expected counts E_k=N/s:
χ² = Σ [(O_k − E_k)^2 / E_k]
degrees of freedom = s − 1
A p-value measures how often a statistic at least this extreme would arise under the null model. It is not the probability that the die is fair. The chi-square approximation uses ample expected counts, including d100. See NIST's method description and SciPy chisquare.
Independence of pairs
A perfectly balanced histogram can conceal a sequence such as 1,1,2,2,3,3,…. We therefore analyze non-overlapping adjacent pairs: (X₁,X₂), (X₃,X₄), …. Each contingency-table dimension has b=min(s,20) categories, with category floor((face−1)·b/s). For d100, five consecutive faces form each category; other dice use one category per face.
Under independence, the expected cell count is E_ij=O_i· O_·j/n_pairs. Pearson's statistic has (b−1)^2 degrees of freedom. We use chi2_contingency(..., correction=False) and require every expected cell count to be at least five. There are 500,000 pairs per batch dataset and 25,000 per fresh-seed dataset. See SciPy chi2_contingency.
This examines categorical dependence in adjacent pairs, not all higher-order patterns. The d100 grouping loses resolution within each bucket. Disjoint pairs avoid treating overlapping observations as independent, but calibration still relies on the IID null model.
One family of 44 tests with Holm adjustment
There are two primary tests per dataset. We adjust all 44 p-values together using Holm at α=0.01. For ordered values p_(1)≤…≤p_(44), the adjusted value at position i is:
p_Holm(i) = min(1, max_{j ≤ i} [(44 − j + 1) · p_(j)])
With valid p-values, this controls family-wise error even when tests are dependent. The significance threshold was not selected after discovering which die looked most irregular. The R statistical documentation explains the procedure and cites Holm (1979).
5. Observed results
Each entry below uses one million faces in Node. These are uniformity p-values before adjustment, so the correction can be checked independently.
| Die | p, xoshiro128ss | p, MT19937 | | --- | ---: | ---: | | d4 | 0.214236 | 0.675186 | | d6 | 0.014352 | 0.335624 | | d8 | 0.914879 | 0.880731 | | d10 | 0.680304 | 0.217666 | | d12 | 0.925451 | 0.122686 | | d20 | 0.187065 | 0.295332 | | d100 | 0.509777 | 0.503956 |
The smallest raw p-value in the entire family was the batch d6/xoshiro result, 0.014352398611…. Its Holm value is 0.631505538885…; the other 43 adjusted values are 1. There were zero rejections at the family threshold of 0.01. An adjusted p-value of 1 follows from correction and capping; it is not a certificate of perfection.
For fresh seeds per face in Node, uniformity p-values were 0.781750 (d4), 0.871192 (d6), 0.669452 (d8), 0.826996 (d10), 0.062821 (d12), 0.779478 (d20) and 0.595757 (d100). Chromium d20 gave χ²=12.2712, 19 degrees of freedom and p=0.873708. All Holm-adjusted independence p-values are also 1. The 22 full records include serial statistics and descriptive effect sizes.
Figure 2 — actually collected xoshiro128ss outcomes. The line marks 5%. The shaded band is a simultaneous binomial null reference envelope, explained below. The vertical axis starts at zero to retain the magnitude of the deviations. Open the SVG.
For batch d20/xoshiro, the mean was 10.50494, versus the model's 10.5; sample variance was 33.19584, versus 33.25. Face 20 appeared 49,487 times and face 12 50,412 times, rather than exactly 50,000 each. The largest absolute deviation was 0.0513 percentage points. The browser maximum deviation was 0.198 percentage points, from a sample twenty times smaller.
Mean and variance are descriptive checks, not two additional primary tests. The records also report empirical total variation distance TV=½ Σ |O_k/N−1/s|, measuring the aggregate observed deviation without automatically calling it population bias.
Figure 3 — comparable residuals. Each bar is (O−Np)/sqrt(Np(1−p)), with p=1/s. Face counts in a histogram are dependent because they sum to N. This is descriptive, not a family of independent normal tests. Open the SVG.
6. Uncertainty: individual intervals and simultaneous envelopes
We publish 95% individual Clopper–Pearson binomial intervals for each face. Batch d20/xoshiro face 1 has frequency 49,682/1,000,000=4.9682% and interval approximately [4.92569%, 5.01097%]. Face 20's individual interval is approximately [4.90627%, 4.99139%], excluding 5%.
This does not contradict the global result. Examining hundreds of individual intervals creates opportunities for apparent exceptions. Individual 95% coverage does not provide 95% simultaneous coverage across all faces. A face selected after collection does not replace the prespecified primary hypothesis.
The plotted bands are instead reference envelopes under H₀, using binomial Bin(N,1/s) quantiles with Bonferroni correction across all 500 face counts in the 22 datasets. Each tail receives 0.01/(2×500). Under the model, the chance that any count falls outside its envelope is at most 1%, accounting for discreteness. These are model-based count references, not simultaneous confidence intervals for unknown face probabilities.
For d20 at N=1,000,000 the envelope is 49,073–50,932 per face; at N=50,000 it is 2,295–2,710. None of the 500 counts lies outside its envelope. This adjustment is separate from Holm for the primary tests. Both constructions are in analyze.py; an exact-interval reference is SciPy's binomial interval API.
7. Linear dependence and its limits
We estimate autocorrelation at lags 1–20:
r(h) = Σ_{t=1}^{N−h} [(X_t − mean)(X_{t+h} − mean)]
/ Σ_{t=1}^{N} (X_t − mean)^2
This is exploratory: 440 coefficients across 22 datasets. The guide uses the normal approximation z_(1−0.01/(2×440))/sqrt(N), with Bonferroni correction. The plot uses the N=50,000 guide, wider than the N=1,000,000 reference. Maximum absolute d20 autocorrelation was 0.001950 for xoshiro batches, 0.007844 for fresh seeds in Node and 0.010067 for Chromium.
Figure 4 — an exploratory diagnostic. Small linear correlation does not exclude nonlinear dependence, other lags or higher-order structure. The guide is asymptotic, not an exact bound. These coefficients are not extra primary tests and do not imply passing TestU01/BigCrush. Open the SVG.
8. Advantage, disadvantage and sums are not uniform
Auditing original faces differs from requiring uniformity after game rules. For independent fair d20 variables X and Y:
P(max(X,Y) = k) = (k/20)^2 − ((k−1)/20)^2 = (2k−1)/400
P(min(X,Y) = k) = ((21−k)/20)^2 − ((20−k)/20)^2 = (41−2k)/400
With advantage, a 20 has probability 39/400=9.75%, while a 1 has probability 1/400=0.25%; disadvantage reverses these. For the sum T of two d6, P(T=t)=(6−|7−t|)/36 for 2≤t≤12; 7 is six times as likely as 2.
Figure 5 — transformations of retained data. Each panel uses 500,000 disjoint pairs from the already-published batch streams. These are not 1.5 million extra fresh rolls or independent tests of every modifier implementation. They illustrate why keep-highest, sums and success-conditioned outcomes require different models. Open the SVG.
9. What deterministic replay verifies
A replay descriptor records algorithm, algorithm version, execution version, mathematical profile, seed material and plan fingerprint. In the mixed API it is nested inside result.replay.rolls. Preserve both formula and descriptor:
import assert from "node:assert/strict";
import { rollMixedDice } from "@erpg/dicecore/systems/mixed";
const original = rollMixedDice("3d20", {
detail: "compact",
randomAlgorithm: "xoshiro128ss",
});
const repeated = rollMixedDice("3d20", {
detail: "compact",
replay: original.replay,
});
assert.deepEqual(
repeated.dice.map((die) => die.rawValue),
original.dice.map((die) => die.rawValue),
);
Every .u8 file stores one raw face per byte. Each .seeds file stores 16 bytes per call, decoded from hexadecimal seedMaterial. The manifest supplies the complete replay template per formula. collect.mjs --verify rebuilds each call and compares every face, as well as checking protocol, engine and stream hashes. This passed for all 14.4 million faces, including the 50,000 browser outcomes.
Replay establishes consistency between inputs and outputs of the identified engine. It does not prove that someone never searched for a favorable seed, that a client was unmodified, or that transmitted history was authentic. Those properties require another protocol, such as prior commitments and authority verification. A published JSON field origin: "crypto" is not itself an independently signed attestation of collection.
10. Reproduction guide
Extract study.zip into an empty directory. Use Node ≥22 and a Python version compatible with the pinned dependencies; the tested environment is recorded above. In that directory:
npm init -y
npm install --ignore-scripts --save-exact @erpg/dicecore@3.7.1
python -m venv .venv
Activate with source .venv/bin/activate on Linux/macOS or .\.venv\Scripts\Activate.ps1 in PowerShell. Alternatively invoke the virtual environment's Python directly.
python -m pip install -r requirements.txt
python verify_package.py --data data --out verified-package.json
python self_test.py
node collect.mjs --verify --out data
python analyze.py --data data --out reproduced-analysis
The first two scripts verify the official package and analysis controls. Replay verification uses published seeds rather than drawing another sample. Analysis independently recalculates all statistics and five figures from raw streams. Compare scientific fields with analysis/statistics.json; floating-point tolerances and runtime metadata can differ across platforms without changing faces.
For an independent experiment, run node collect.mjs --out new-data. Serve local files with python -m http.server 8080 --bind 127.0.0.1, open http://127.0.0.1:8080/browser.html, download its three files and import them with node collect.mjs --import-browser browser-dataset.json --out new-data. Analysis expects all 22 datasets. New seeds yield new counts: retain adverse results and never retry until a study passes. The README explains paths, formats, checksums and browser bundle rebuilding.
Positive controls include balanced counts, artificial bias and balanced marginals with repeated pairs. The actual study analysis detects both artificial defects. Holm matches a hand-computed example and intervals match SciPy's exact implementation. The control script is published so a favorable finding does not depend on an analysis incapable of detecting defects.
11. Library, source, licensing and integrity
The measured dependency is npm package 3.7.1. That artifact declares MIT and includes the license preserved in this study. The public fork repository exposes source for inspection, but its current license has separate Arkanus terms. Publicly readable source does not automatically make every release and branch open source on identical terms. Inspect the license of the artifact and revision you use. The original upstream library is also available.
We inspected source revision 1940044c34c47a3c55faa723ccb0d524eef79b03. Package metadata does not identify gitHead, so we do not claim that this commit produced the published tarball. The executable identity is established through package bytes, integrity and manifest fingerprints, rather than an assumed version-to-commit correspondence.
Package: @erpg/dicecore@3.7.1
Tarball SHA-512 SRI:
sha512-3Aa8dkiJLwXOMGktEj3lZXJtHqYLaGUSeRLuFKqaP795lVmSYPW9gjY26IY8LTiWXSmTbFvBjEa9IlBWRQbUCg==
Tarball SHA-256:
d4defb2ff92b9bf56cb3ed2ba9cee276d18d2ba9ae4ba2f09d2b08c013d368a9
The package verification report, evidence and scope record and download checksum list complete the trail. Hashes verify consistency of bytes; they do not replace signatures or independent audits.
12. The boundaries of the conclusion
This work supplies three verifiable pieces of evidence: a mapping that removes modulo bias under its assumptions; observed faces compatible with the defined statistical models; and complete replay through the identified engine. Together these support auditable use of the common dice examined.
The limits are concrete: seven dice types, 22 profiles, two PRNGs, one browser and one Node environment; limited pair and correlation diagnostics; chi-square and normal approximations; no power analysis for every possible alternative. Failure to detect bias does not establish a universal upper bound on all possible bias. An implementation detail is that xoshiro's all-zero 128-bit state receives a deterministic repair; the collected histogram does not prove behavior over all states.
This does not certify cryptographic security, client tamper resistance, the Go runtime, every modifier, every platform or the package currently served to every visitor. Changes to package, algorithm or seeding require identifying the new bytes and repeating the assessment. Cite version, protocol and results while retaining these boundaries.
Try the dice roller or read Why Fair Dice Feel Rigged for the experience of streaks under a fair model. To verify this conclusion, use the published files: reproduction relies on the data and scripts rather than trust in the article.