Credibility and measurement

Results with period, baseline and limits.

Not every workflow needs the same metric. The site separates predictive performance, operating reliability and safety controls.

General principle

Limitations of rules, indicators and models

Outputs produced by rules, indicators and models are experimental or informational results, not decisions.

They may be incomplete, unstable or incorrect, especially when data are missing or delayed, exceptional events occur, regimes change, or conditions differ from those observed during design and validation.

Where relevant and available, each result should be interpreted together with its baseline, observation period, sample size, metrics, measure of uncertainty and known failure conditions.

Results and evidence status

What is measured, under evaluation or operational.

The table does not invent a common score: it separates predictive evidence, operational reliability and activities still maturing.

WorkflowEvidence statusMaturityDetail
WF-1 · Weather automationOperationalActive · educationalOpen the project page for available metrics, period and limitations.
WF-2 · The New MarginUnder evaluationActive · experimentalOpen the project page for available metrics, period and limitations.
WF-3 · Market OverviewUnder evaluationActive · experimentalOpen the project page for available metrics, period and limitations.
WF-4 · AI Supply ChainUnder evaluationActive · experimentalOpen the project page for available metrics, period and limitations.
WF-5 · SatelliteOperationalActive · event-conditionedOpen the project page for available metrics, period and limitations.
WF-6 · Energy Crisis ThermometerUnder evaluationActive · experimentalOpen the project page for available metrics, period and limitations.
WF-7 · Etna ForecastMeasured in workflowActive · educationalOpen the project page for available metrics, period and limitations.
WF-8 · Bluesky multi-sourceOperationalActive · experimentalOpen the project page for available metrics, period and limitations.
WF-9 · World Economy EngineUnder evaluationActive · experimental · evidence-awareOpen the project page for available metrics, period and limitations.
WF-10 · Etna SentinelOperationalActive · educational · fail-closedOpen the project page for available metrics, period and limitations.
WF-11 · Hydrogen Route ObservatoryInspectable calculationExperimental · calculation · evidence-awareOpen the project page for available metrics, period and limitations.
WF-12 · Ammonia & Fertilizer ChainForecast maturingExperimental · calculation + forecastOpen the project page for available metrics, period and limitations.
WF-13 · Italy Variable Renewable ForecastObserved source conditionalExperimental · forecast · point-in-timeOpen the project page for available metrics, period and limitations.
WF-14 · Process SentinelSynthetic benchmarkExperimental · simulated plantOpen the project page for available metrics, period and limitations.
WF-15 · Virtual AnalyzerSynthetic benchmarkExperimental · simulated laboratoryOpen the project page for available metrics, period and limitations.
WF-16 · Energy & Steam OptimizerSimulated optimizationExperimental · hybrid optimizationOpen the project page for available metrics, period and limitations.

Measured in workflow means that the project page publishes specific metrics; under evaluation means relevant metrics are not yet centralised; operational concerns checks, runs and traceability rather than predictive skill.

Available general metrics

Forecasts

Error, bias, coverage, directional accuracy and baseline comparison, only on matured outcomes.

Classifications and anomalies

Precision, recall or anomaly frequency when reliable labels exist; otherwise context and human review.

Operating workflows

Successful, failed, skipped and blocked runs, missing data and prevented duplicates.

Transparency rule

Never show only the best metric; state period, sample size and preliminary status.

Workflow status

Workflow status
WorkflowStatusCadenceRelevant metrics
WF-1 · Weather and neck-discomfort indexActive · educationalMonday, Wednesday and FridayRun continuity, missing data caught and duplicates prevented. No clinical performance metric is claimed.
WF-2 · Refining margins and forecastActive · experimentalWeeklyMAE, median error, bias and interval coverage, always with period and matured sample size.
WF-3 · Market OverviewActive · experimentalWeekdaysDirectional accuracy, average move, errors, period stability and baseline comparison.
WF-4 · AI Supply ChainActive · experimentalWeeklyModel-versus-naïve comparison, outlook error and basket stability over time.
WF-5 · Satellite — event-based EtnaActive · event-drivenChecked every 6 hours; posted only for a valid new eventRun outcomes (published, skipped, blocked), source availability and duplicates prevented. No operational dispersion skill is claimed.
WF-6 · Energy Crisis ThermometerActive · experimentalSeveral weekly runsDriver coverage, shipping confidence, OOS forecast metrics, Dynamic versus Static, drawdown and period stability.
WF-7 · Etna ForecastActive · educationalMonday, Thursday, Sunday and eventsBrier score, comparison with climatology/persistence/Hawkes, calibration, matured sample and gate-selected source.
WF-8 · Bluesky multi-sourceActive · experimentalMonday, Wednesday and FridaySource-gate audit, publication receipts, grapheme count, multi-image output and PDS/AppView verification.

Status describes project maturity and does not guarantee continuous availability of external sources.

How to use this evidence

Different metrics answer different questions.

Technical reviewers

Delivery capability

Inspect testing, modularity, logging, schedules, error handling and communication of limitations.

Professional evidence →
Data not yet centralised: the site does not invent an overall run-success rate. It will be shown only when a shared automated register supplies it.