Credibility and measurement

Results with period, baseline and limits.

Not every workflow needs the same metric. The site separates predictive performance, operating reliability and safety controls.

General principle

Limitations of rules, indicators and models

Outputs produced by rules, indicators and models are experimental or informational results, not decisions.

They may be incomplete, unstable or incorrect, especially when data are missing or delayed, exceptional events occur, regimes change, or conditions differ from those observed during design and validation.

Where relevant and available, each result should be interpreted together with its baseline, observation period, sample size, metrics, measure of uncertainty and known failure conditions.

Results and evidence status

What is measured, under evaluation or operational.

The table does not invent a common score: it separates predictive evidence, operational reliability and activities still maturing.

WorkflowEvidence statusMaturityDetail
WF-1 · Weather automationOperationalActive · educationalOpen the project page for available metrics, period and limitations.
WF-2 · The New MarginUnder evaluationActive · experimentalOpen the project page for available metrics, period and limitations.
WF-3 · Market OverviewUnder evaluationActive · experimentalOpen the project page for available metrics, period and limitations.
WF-4 · AI Supply ChainUnder evaluationActive · experimentalOpen the project page for available metrics, period and limitations.
WF-5 · SatelliteOperationalActive · event-conditionedOpen the project page for available metrics, period and limitations.
WF-6 · Energy Crisis ThermometerUnder evaluationActive · experimentalOpen the project page for available metrics, period and limitations.
WF-7 · Etna ForecastMeasured in workflowActive · educationalOpen the project page for available metrics, period and limitations.
WF-8 · Bluesky multi-sourceOperationalActive · experimentalOpen the project page for available metrics, period and limitations.
WF-9 · World Economy EngineUnder evaluationActive · experimental · evidence-awareOpen the project page for available metrics, period and limitations.
WF-10 · Etna SentinelOperationalActive · educational · fail-closedOpen the project page for available metrics, period and limitations.
WF-11 · Hydrogen Route ObservatoryUnder evaluationExperimental · calculation · evidence-awareOpen the project page for available metrics, period and limitations.
WF-12 · Ammonia & Fertilizer ChainUnder evaluationExperimental · calculation + forecastOpen the project page for available metrics, period and limitations.
WF-13 · Italy Variable Renewable ForecastUnder evaluationExperimental · forecast · point-in-timeOpen the project page for available metrics, period and limitations.
WF-14 · Process SentinelControlled benchmarkExperimental · synthetic data (simulated data) · anomaly detectionOpen the project page for available metrics, period and limitations.
WF-15 · Virtual AnalyzerControlled benchmarkExperimental · synthetic data (simulated data) · soft sensorOpen the project page for available metrics, period and limitations.
WF-16 · Energy & Steam OptimizerControlled benchmarkExperimental · hybrid public + synthetic data (simulated data) · optimisationOpen the project page for available metrics, period and limitations.
WF-17 · Etna Fusion LabProspective maturationActive · experimental · quality-awareOpen the project page for available metrics, period and limitations.
WF-18 · Climate Comfort HoursOperational · heuristicActive · educational · heuristicOpen the project page for available metrics, period and limitations.
WF-19 · Lab Intelligence CommunityOperational · editorialActive · editorial · automatedOpen the project page for available metrics, period and limitations.
WF-20 · The Public Service RunUnder evaluation · calibrationActive · educational · experimentalOpen the project page for available metrics, period and limitations.
WF-21 · Port Operability · NYC FerryProspective maturationActive · research shadow · prospective verificationOpen the project page for available metrics, period and limitations.

Measured in workflow means that the project page publishes specific metrics; under evaluation means relevant metrics are not yet centralised; operational concerns checks, runs and traceability rather than predictive skill.

Available general metrics

Forecasts

Error, bias, coverage, directional accuracy and baseline comparison, only on matured outcomes.

Classifications and anomalies

Precision, recall or anomaly frequency when reliable labels exist; otherwise context and human review.

Operating workflows

Successful, failed, skipped and blocked runs, missing data and prevented duplicates.

Transparency rule

Never show only the best metric; state period, sample size and preliminary status.

Workflow status

Workflow status
WorkflowStatusCadenceRelevant metrics
WF-1 · Weather automationActive · educationalMonday, Wednesday and FridayRun continuity, missing data caught and duplicates prevented. No clinical performance metric is claimed.
WF-2 · The New MarginActive · experimentalWeeklyMAE, median error, bias and interval coverage, always with period and matured sample size.
WF-3 · Market OverviewActive · experimentalWeekdaysDirectional accuracy, average move, errors, period stability and baseline comparison.
WF-4 · AI Supply ChainActive · experimentalWeeklyModel-versus-naïve comparison, outlook error and basket stability over time.
WF-5 · SatelliteActive · event-conditionedChecked every 6 hours; post on a valid new eventRun outcomes (published, skipped, blocked), source availability and duplicates prevented. No operational dispersion skill is claimed.
WF-6 · Energy Crisis ThermometerActive · experimentalSeveral weekly runsDriver coverage, shipping confidence, OOS forecast metrics, Dynamic versus Static, drawdown and period stability.
WF-7 · Etna ForecastActive · educationalMonday, Thursday, Sunday and eventsBrier score, comparison with climatology/persistence/Hawkes, calibration, matured sample and gate-selected source.
WF-8 · Bluesky multi-sourceActive · experimentalMonday, Wednesday and FridaySource-gate audit, publication receipts, grapheme count, multi-image output and PDS/AppView verification.
WF-9 · World Economy EngineActive · experimental · evidence-awareWeekly · MondayExpanding backtest with purge gap, prequential selection, naive benchmark, bootstrap skill, 90% conformal intervals, graph ablation and anti-leakage/vintage audit.
WF-10 · Etna SentinelActive · educational · fail-closedTue/Thu/Sun + VONA monitor every 3 hoursUnit/end-to-end tests, data-quality and disclosure gates, deduplication/idempotency, fail-closed on uncertain/positive official status and immutable history.
WF-11 · Hydrogen Route ObservatoryExperimental · calculation · evidence-awareMon–Fri 10:50 Europe/RomePhysical/dimensional invariants, sourcing coherence, lineage, snapshot replay.
WF-12 · Ammonia & Fertilizer ChainExperimental · calculation + forecastMon–Fri 09:50 Europe/Rome; monthly maturationWalk-forward, persistence, seasonality, gas-only and physical baselines; MAE/RMSE/bias/coverage/skill.
WF-13 · Italy Variable Renewable ForecastExperimental · forecast · point-in-timeDaily 08:50 Europe/RomeTime split and persistence; proxy separated from actuals and never matured as a real forecast; prospective ENTSO-E A69 baseline when available; nMAE/RMSE/bias/coverage/skill.
WF-14 · Process SentinelExperimental · synthetic data (simulated data) · anomaly detectionEvery 6 h: 00:50/06:50/12:50/18:50 UTCParameters frozen on a separate development simulation; same frozen 28-day synthetic data (simulated data) benchmark; event recall, missed rate, false-alarm episodes/24h, P50/P90 delay, v1.4.16 ensemble comparison and anti-leakage.
WF-15 · Virtual AnalyzerExperimental · synthetic data (simulated data) · soft sensorEvery 6 h: 01:50/07:50/13:50/19:50 UTCRMSE/MAE/bias/coverage, skill versus last lab, split-conformal and truth-vs-lab; risk-coverage diagnostic only.
WF-16 · Energy & Steam OptimizerExperimental · hybrid public + synthetic data (simulated data) · optimisationDaily 11:50 Europe/RomeFeasibility, balances, optimality, solve time, stability and reference policy.
WF-17 · Etna Fusion LabActive · experimental · quality-awareHourly nowcast · social Tue/Thu/SatImmutable issue ledger, reconciliation sidecars, prospective calibration, hysteresis/cooldown and adaptive learner kept in shadow before gates.
WF-18 · Climate Comfort HoursActive · educational · heuristicTue/Thu 17:00 · Sun 21:00 Europe/RomeIndependent city/day validation, batch→recent cache→single retry, grey cells for missing data, publication only with 4 core cities and at least 80% of the basket.
WF-19 · Lab Intelligence CommunityActive · editorial · automatedMon/Wed/Fri/Sun fixed · Tue/Sat FLEXslot and blackout validation, weekly cap, deduplication, reviewable diff and fail-closed gates for schedule, policy and credentials
WF-20 · The Public Service RunActive · educational · experimentalDaily forecast 05:17 · social Tue IT / Fri ENMAE, RMSE, MAPE, median/p90 absolute error, P10–P90 coverage, top-1 departure accuracy, departure regret and explicit warm-up stage
WF-21 · Port Operability · NYC FerryActive · research shadow · prospective verificationDaily run + 30-min event watcher · social Fri IT / Mon ENImmutable forecasts verified at D+1 and D+7; prospective metrics include episode recall, lead time, alert burden on verified available days, horizon coverage and restricted calibration; UNKNOWN stays outside denominators.

Status describes project maturity and does not guarantee continuous availability of external sources.

How to use this evidence

Different metrics answer different questions.

Technical reviewers

Delivery capability

Inspect testing, modularity, logging, schedules, error handling and communication of limitations.

Professional evidence →
Data not yet centralised: the site does not invent an overall run-success rate. It will be shown only when a shared automated register supplies it.