Approach for providing data series required in prognostic simulations
Viesturs Berzins, Uldis Bethers, Klaus-Peter Holz, Juris Sennikovs
Prognostic applications of mathematical models of hydraulic processes require input data sets of meteorological parameters. These data sets referred hereinafter as forcing data should be rankable via quantitative parameters in order to calculate “typical” or “critical” scenarions. The providing of adequate typical / critical forcing data is common problem for wide range of applications: prediction of morphodynamic response, siltation, oil spills, environmental impact assessment, forecast of ecosystem trends. An approach of generation of the nearly – natural artificial data sets is proposed in this paper. Approach is aimed in replacing the commonly used practice of employment of historic forcing data series for the prognostic calculations.
The mathematical modelling of almost any natural process is highly sensitive with respect to the model input data. The forcing data of hydraulic systems usually is meteorological data. The response of the model to the forcing data is as a rule non-linear. Therefore the quality of the input data is of key importance. The impact of the model input data series on the modelling results is comparable to the impact of the overall quality of the model formulations. However, usually modelers pay much more attention to the appropriate selection of the model formulation, modelling software and domain geometry, leaving the forcing data selection and quality issues off the primary priority list.
The non-steady-state modelling tasks may be classified according to the
necessary forcing input data sets as
1.
Hindcasts. Historical data sets are required and existing observation series may be
used as a rule; data quality is the most crucial issue.
2.
Operational (or real-time)
modelling. The actual observations should be provided
for the modelling system to ensure nearly on-line simulations. The main issues
are in organising real-time data flow, generation of short-time forecast of
input data, and observation data assimilation.
3.
Forecasts (prognostic calculations). As a rule forecasts are used to predict the
behaviour of the system under consideration after some alterations are
introduced. Usually, evaluation of the typical and critical behaviour
(scenarious) is required. The critical issue here is in providing the
appropriate forcing data series, which may be regarded as typical or critical
for predictions.
This paper is aimed in proposing approach for building of forcing data sets for the prognostic calculations. We do not consider problems associated with data quality, data assimilation, and organising nearly-real-time information flows.
The approach is developed for providing the prognostic time-series for the wind data required as input data for the modelling of littoral drift along particular stretches of Latvian and German coastline. However, we have kept the description of the proposed method as general as possible, giving the particular concretisation from that application only where necessary.
The proposed data-series generation method consists of two parts:
1.
Analytic part analyses the
historical data series, extracting the rules of its structure.
2.
Synthetic part uses these rules
to create artificial nearly – natural data series with prescribed properties
(severeness).
The following stepwise scheme is proposed for the derivation of the main
[statistical] rules of the meteorological data structure (see also fig. 1)
1.
Expert selection of event, a
quasi-independent entity containing sub-series of meteorological data.
Preferably, the events should be (1) meaningful by themselves, (2) as
independent of the meteorological data in preceding events as possible.
2.
Development of the set of
qualitative and quantitative descriptors for the events, i.e. event classification.
3.
Event recognition in the
historic data series, i.e. splitting of the data series into events. This may
be done via
3.1.
Expert (“manual”) splitting
of the part of the available historic data series into events. The formal rules developed under p.2 together with expert’s knowledge is
used in this process; the set of rules may be complemented by new ones.
3.2.
Development of automated
method for the event recognition in the meteorological data series. The
possible options are
3.2.1.
Direct algorithm, employing
the qualitative descriptors p.2 for the event recognition. The “training” of
the direct algorithm is assumed by means of introducing new descriptors and
rules, until the performance of the algorithm is comparable with the expert
recognition of p.3.1.
3.2.2.
AI (artificial intelligence)
algorithm, assuming the building of self-teaching neural network, which is being
trained by the results of p.3.1. The descriptors of p.2 have been employed in
this case; the improvement of the algorithm is possible via extension of either
the neural network complexity or training process.
3.3.
Automated event recognition
in the whole available meteorological data set.
4.
Analysing the rules
and statistics of event occurrence. The probabilities of event
occurrence along with the statistics of their sequence are extracted at this
step from the results of p.3.3.
5.
Analysing the strength
and relevance of the events. The step may be splitted into
5.1.
Quantification of the strength
(or severeness) of the events in meteorological sense. The association of the
events with some “measurable” integral parameter is necessary for further
ranking of the sets of events, and development of criterions describing typical
or critical conditions. However, meteorological conditions may be
self-contradictory when applying qualitative judgements about their severeness.
Example: “mild winters (regarding temperature) are stormy (due high occurrence
of cyclones) in

5.2.
Mapping the meteorological events on ANY particular process under consideration
(for example, integral load transport). Mapping tool would be simple
mathematical model (eg. formulae). Mapping ensures further evaluation of event relevance.
5.3.
Quantification of the relevance
of the events in the sense of the response (for instance, morphodynamical) of
the system to the meteorological forcing.
5.4.
Analysing the rules
and statistics of event strength and relevance, ranking of events
according to these parameters.
6.
Analysing the rules
and statistics of the data sequences inside the events. This step is
aimed for finding
6.1.
Rules (including
probabilistic) of temporal sequences of particular meteorological parameters.
6.2.
Correlations and/or functional relationships
between different meteorological parameters.
The above pp. 1 to 6 are preparatory steps, i.e. meteorological data analysis. Although the analytic part of the proposed method seems rather expanded, one may consider that it must be employed only once.
After the analysis of the meteorological data we propose the forcing
data generation scheme in three basic steps, see also fig. 2.
7.
Generation of the artificial sequence of events. The event
sequence is generated basing on the previously determined rules about (1) the
sequence of events, i.e. probability distribution of transfer from one event
type to another, and (2) the probabilities of the event duration. The average
probabilities may be used to generate the typical sequence of events. The
probability distributions may be modified (weighted) to obtain different (say,
critical) artificial data sets (primary moderation). Thus, for critical data
sets one may (1) increase the probabilities of
transition to more relevant (in sense of system response, as aquired by process
mapping) events; (2) increase the probabilities that these relevant events last
longer.
8.
Generation of the artificial sequences of [primary] meteorological
data inside the particular events. The data sequences are generated
from the previously determined rules about the probability distributions of the
particular quantitative parameters defining the meteorological data sequence
inside the event. The data items forming the actual meteorological conditions
are not independent; therefore the generation of the data sequence inside the
event should be done for selected [primary] parameter: either most relevant
parameter for further prognostic modelling, or parameter clearly determining
other meteorological data. The probability distributions of the quantitative
descriptors may be modified to increase/decrease the severeness and/or
relevance of the particular event (secondary moderation).
9.
The building of the
artificial sequence of the meteorological (or forcing, or model input) data by
filling the sequences of p.8 into events of p.7.
Thus, sequence pp. 7 to 9 ensures the generation of nearly – natural forcing data set. The data set is typical if the average probabilities are used where required. The data set may be turned more/less critical by both/either (1) changing the occurrence/duration of events with different relevance, (2) changing the probabilistic parameters, determining the sequence of the data within particular event.

The proposed approach was applied for providing artificial data sets for
longshore load transport modelling (Grzibovskis et. al., 1999). The time-series of the wind measurements
(velocity and direction) was primarily required for numerical simulations. Far
sea wave data was calculated by fetch model (Sennikovs
et. al., 1998), whilst water level considered as less significant, and found
from the wind parameters via correlation analysis. Seasonal, annual and
long-term data series were required.
Continuous data series for six years (LHMA, 1992-97) of almost full set
of meteorological observations at Riga station (air temperature, humidity,
pressure, wind speed and direction, total and high cloudiness, precipitation),
synoptic pressure charts (LHMA, 1985-97) for thirteen years, and wind data
since 1971 were available for the analysis. Meteorological observations were
recorded 8 times daily, whilst pressure charts were issued 5 times per week (on
workdays).
The primary data required for morphodynamic
calculations is wind data. However, we failed to identify sub-sets of wind data
to qualify as events. The criterions applied for event selection were (1) an
event must be simple but meaningful, (2) contents of the weather conditions
must follow some general rules within an event, (3) duration of an event must
be in synoptic time scale, (4) weather conditions during the actual event must
be as independent as possible of preceding events, i.e. development of weather
conditions should be self-containing within an event. All above criterions hold
true for the time periods, associated with the passage of [particular periphery
of] some baric formation (cyclon L, or anticyclon H) over the place of observation.
The classification of events
was assumed according to the type of baric formation (cyclon L, anticyclon H, or neutral formations N) and the periphery
passing over the observation place (8 peripheries for L and H). The recognition
of the baric type was assumed possible from the air pressure, whilst periphery
– from the wind time-series. Region of genesis (Icelandic, South European,
local for L; Azoric, Greenlandic, East European,
South European or local for H) was also considered as classification criteria.
However, the recognition of the region of genesis is possible mainly via
analysis of a sequence of pressure charts as well as from the seasonal
temperature and humidity characteristics. The automated analysis of graphical
information was considered too complicated, whilst temperature and humidity was
unrelevant for the particular application.
The air pressure was assumed
as the basic criteria for recognising of the baric
formation. Formation was assumed as L, if
The expert (“manual”)
recognition of the events were performed for the whole
6 years time period, identifying also the region of genesis. 31,7% of total time may be associated with cyclons,
45,5% with anticyclons, whilst 22,8% with neutral
conditions. Most (68%) of cyclons were Icelandic,
generally moving from W to E. These cyclons mainly
passed the Northern part of
The direct algorithm of the
recognition of the baric formations was based on the same air pressure criteria
as expert recognition. We assumed L for
·
dividing transient periods between preceding and
following events instead of assuming them neutral if transition time is less
than 24 h;
·
eliminating weak cyclons (
The periphery of the baric
formation was detected according to the following algorithm:
·
the average wind direction for the beginning (12 h),
end (12 h) and whole event was calculated;
·
the periphery of the cyclons
was detected by subtracting 60°, whilst
periphery of anticyclons by adding 150° from/to the average wind
direction during the events;
·
the turning angle a of the wind was used for
the detection of whether the center of baric formation passes the place of
observation. We voluntary supposed that it is the case, if ½a½ > 135°. The comparison of the
expert and automated event recognition is given in table 1.
Table 1 Comparison of expert and automated
event recognition
|
|
H |
L |
N |
|||
|
Expert |
Automatic |
Expert |
Automatic |
Expert |
Automatic |
|
|
General fit, % |
- |
97,2 |
- |
98,6 |
- |
83,7 |
|
Average duration, h |
98,0 |
131,0 |
81,5 |
79,6 |
59,0 |
71,9 |
|
Periphery accuracy 90 ° and better, % |
- |
91,7 |
- |
95,6 |
- |
- |
Generally, most (97 to 99 % of H/L) of observations were found as belonging to the “right” baric formations. Disagreement between the expert and automatic event recognition is associated with weakly-defined formations. The comparison of average event duration indicated that direct algorithm fails to recognize sequential passing of several anticyclons, when air pressure remains high also during the transition period. The typical wind speeds during these anticyclons are rather low, therefore the wind data analysis does not improve situation; on the other hand the expected relevance of the anticyclonic events is low.
The “accuracy” of the determining of periphery is about 90° by expert evaluation, whilst 45° was used for the automated periphery recognition. One may consider the performance of direct algorithm as rather good. Note, that more distinct (short and strong) cyclonic events are processed better.
The attempts of applying AI approach for event recognition lead to less reasonable results, and they are omitted in this paper.
The morphodynamic
relevance of different events was assessed by mapping them to the integral longstore load transport through single cross shore profile
at Latvian coast (orientation from SSW to NNE). The average load transport (per
hour) in both direction during 17 types of events is summarised in table 2.
One
may conclude that (1) the same events may be responsible for sediment movement
in both directions due wind turns during the event, (2) anticyclons
have lower morphodynamic relevance as cyclons, (3) only two to three peripheries of both cyclons and anticyclons (i. e. ~
30% of events) may be of morphodynamical
significance.
Table 2 Morphodynamic
relevance of different events
|
Event
type |
QN,
m3/h |
Qs,
m3/h |
Q,
m3/h |
|
L-N L-NE L-E L-SE L-S L-SW L-W L-NW |
1 0 4 17 123 67 2 1 |
-4 -1 -2 -5 -47 -321 -88 -27 |
3 -1 2 12 76 -254 -86 -26 |
|
H-N H-NE H-E H-SE H-S H-SW H-W H-NW |
3 10 4 1 0 0 1 1 |
-10 -32 -34 -38 -8 -2 -6 -2 |
-7 -22 -30 -37 -8 -2 -5 -1 |
|
N |
4 |
-9 |
-5 |
Two basic probability distribution
groups (probability table of transition between different events, and
probabilities of duration of different events) were calculated, and applied for
obtaining random (but nearly-natural) event sequence. Both probability
distributions are essentialy seasonal.
The principal modelling input data is wind speed and direction. Therefore
the analysis of the internal rules of the wind data inside the events was
performed. Principal simplification was applied assuming that wind speed and
direction may be considered independently.
The ability of recognition
of the event from changing wind direction (turning) was demonstrated before.
Thus, we applied the general rule of changing wind direction clockwise or
anticlockwise during the passage of baric formation. The approach of describing
wind direction time development by three angles (a1 at the start
moment, a2 at the middle, a3 at the end
moment of event) was used. 17 ´ 3 probability distributions
, were calculated for all event types ´´.
The following general rules
of changing wind velocity during the event were applied:
1.
The wind speed in the center of H is lower as in
periphery (see typical example in fig. 3).
2.
The wind speed in the center and peripheries of cyclons is lower as in some [circular] intermediate zone
(see example in fig. 4).
3.
The wind speed is rather low during neutral events,
whilst it is either falling or raising during
transition between events.
First two rules allow to distinguish three principally different time-developments
of wind speed within an event:
1.
Wind speed has one minimum during the passage of the
anticyclon (as in fig. 3).
2.
Wind speed has two maximums and central minimum, if
the central part of the cyclon passes close to the
observation point (as in fig. 4).
3.
Wind speed has one maximum during the passage of the
cyclon periphery.

Fig.3 Wind speed (m/s) during passage of
(eastern periphery) of anticyclon, May, 1992; time in
days
We applied the following
scheme of the generation of the wind speed sequences. Analytic part (time and
wind speed normalized to event duration and maximum wind speed):
1.
Approximation of the wind speed time development by
4th order polynomial expression for all recognized events by least
square method.
2.
Finding 5 characteristic nodes (ti,
Wi), i = 1, 2… 5
for each event: (1) wind speed at the beginning and end of events, (2) wind
speeds and their occurrence times at wind speed extremes N £ 3, (3) 5 – N – 2 wind speed
and time values in intermediate time moments.
3.
Calculation of 8 probability distributions for each
event type:
, i = 2, 3, 4;
, i = 1, 2… 5 (´´ is event type).
4.
Calculation of the probability distributions for
maximum dimensional wind speed
and the
non-dimensional deviation of the maximum of approximation curve from the
observed wind speed maximum
.
Synthetic part, i.e.
generation of wind speed time development during the particular event was
performed by the following algorithm:
5.
Random determination of the event’s strength Wmax from the probability distribution
.
6.
Random determination of the non-dimensional time
moments t2, t3, t4 and non-dimensional wind
speeds W1, W2, W3, W4, W5
from the probability distributions
, i = 2, 3, 4 and
, i = 1, 2… 5.
7.
Construction of polynomial interpolation curve through
five characteristic points (ti, Wi).
8.
Random perturbation of the discrete analogue of the
interpolation curve according to probability distribution
.
9.
Transition to dimensional time and wind speed.

Fig.4 Wind
speed (m/s) during passage of center and SW periphery
of cyclon, Feb-1992; time in days
Above approach resulted in artificial whilst nearly natural time-series of wind data, further employed for the prognostic morphodynamic calculations.
The proposed approach to the knowledge of authors is rather innovative and may be widely employed for providing input data required for a range of different prognostic calculations.
At the same time (1)
particular steps of the proposed method should be carefully tested and improved
as well as (2) procedures for the quality assessment of the generated data sets
should be developed and employed. Both lather issues regretfully fall outside
the scope of this paper.
References
R.Gržibovskis,
U.Bethers, J.Sennikovs 1999. “Modelling littoral drift along Latvian
and
J. Senņikovs, R. Gržibovskis,
U.Bethers, K.-P. Holz 1998. Multi-level approach to the estimation of the load transport near Ruegen. Proceedings ICHE'98,
LHMA 1992-97. Tables TMS. Technical reports. Hydrometeocenter of the State Hydrometeorological Agency of Latvia. In Latvian.
LHMA 1985-97. Meteorological bulletins of the Latvian Hydrometeorological Agency. Hydrometeocenter of the State Hydrometeorological Agency of Latvia. In Latvian.