Approach for providing data series required in prognostic simulations

Viesturs Berzins, Uldis Bethers, Klaus-Peter Holz, Juris Sennikovs

 

Prognostic applications of mathematical models of hydraulic processes require input data sets of meteorological parameters. These data sets referred hereinafter as forcing data should be rankable via quantitative parameters in order to calculate “typical” or “critical” scenarions. The providing of adequate typical / critical forcing data is common problem for wide range of applications: prediction of morphodynamic response, siltation, oil spills, environmental impact assessment, forecast of ecosystem trends. An approach of generation of the nearly – natural artificial data sets is proposed in this paper. Approach is aimed in replacing the commonly used practice of employment of historic forcing data series for the prognostic calculations.

 

 

Introduction

The mathematical modelling of almost any natural process is highly sensitive with respect to the model input data. The forcing data of hydraulic systems usually is meteorological data. The response of the model to the forcing data is as a rule non-linear. Therefore the quality of the input data is of key importance. The impact of the model input data series on the modelling results is comparable to the impact of the overall quality of the model formulations. However, usually modelers pay much more attention to the appropriate selection of the model formulation, modelling software and domain geometry, leaving the forcing data selection and quality issues off the primary priority list.

 

The non-steady-state modelling tasks may be classified according to the necessary forcing input data sets as

1.      Hindcasts. Historical data sets are required and existing observation series may be used as a rule; data quality is the most crucial issue.

2.      Operational (or real-time) modelling. The actual observations should be provided for the modelling system to ensure nearly on-line simulations. The main issues are in organising real-time data flow, generation of short-time forecast of input data, and observation data assimilation.

3.      Forecasts (prognostic calculations). As a rule forecasts are used to predict the behaviour of the system under consideration after some alterations are introduced. Usually, evaluation of the typical and critical behaviour (scenarious) is required. The critical issue here is in providing the appropriate forcing data series, which may be regarded as typical or critical for predictions.

This paper is aimed in proposing approach for building of forcing data sets for the prognostic calculations. We do not consider problems associated with data quality, data assimilation, and organising nearly-real-time information flows.

 

The approach is developed for providing the prognostic time-series for the wind data required as input data for the modelling of littoral drift along particular stretches of Latvian and German coastline. However, we have kept the description of the proposed method as general as possible, giving the particular concretisation from that application only where necessary.

 

 

Formal scheme

The proposed data-series generation method consists of two parts:

1.      Analytic part analyses the historical data series, extracting the rules of its structure.

2.      Synthetic part uses these rules to create artificial nearly – natural data series with prescribed properties (severeness).

 

Analytic sequence of operations

The following stepwise scheme is proposed for the derivation of the main [statistical] rules of the meteorological data structure (see also fig. 1)

 

1.      Expert selection of event, a quasi-independent entity containing sub-series of meteorological data. Preferably, the events should be (1) meaningful by themselves, (2) as independent of the meteorological data in preceding events as possible.

2.      Development of the set of qualitative and quantitative descriptors for the events, i.e. event classification.

3.      Event recognition in the historic data series, i.e. splitting of the data series into events. This may be done via

3.1.   Expert (“manual”) splitting of the part of the available historic data series into events. The formal rules developed under p.2 together with expert’s knowledge is used in this process; the set of rules may be complemented by new ones.

3.2.   Development of automated method for the event recognition in the meteorological data series. The possible options are

3.2.1.      Direct algorithm, employing the qualitative descriptors p.2 for the event recognition. The “training” of the direct algorithm is assumed by means of introducing new descriptors and rules, until the performance of the algorithm is comparable with the expert recognition of p.3.1.

3.2.2.      AI (artificial intelligence) algorithm, assuming the building of self-teaching neural network, which is being trained by the results of p.3.1. The descriptors of p.2 have been employed in this case; the improvement of the algorithm is possible via extension of either the neural network complexity or training process.

3.3.   Automated event recognition in the whole available meteorological data set.

4.      Analysing the rules and statistics of event occurrence. The probabilities of event occurrence along with the statistics of their sequence are extracted at this step from the results of p.3.3.

5.      Analysing the strength and relevance of the events. The step may be splitted into

5.1.   Quantification of the strength (or severeness) of the events in meteorological sense. The association of the events with some “measurable” integral parameter is necessary for further ranking of the sets of events, and development of criterions describing typical or critical conditions. However, meteorological conditions may be self-contradictory when applying qualitative judgements about their severeness. Example: “mild winters (regarding temperature) are stormy (due high occurrence of cyclones) in Latvia”.

 



5.2.   Mapping the meteorological events on ANY particular process under consideration (for example, integral load transport). Mapping tool would be simple mathematical model (eg. formulae). Mapping ensures further evaluation of event relevance.

5.3.   Quantification of the relevance of the events in the sense of the response (for instance, morphodynamical) of the system to the meteorological forcing.

5.4.   Analysing the rules and statistics of event strength and relevance, ranking of events according to these parameters.

6.      Analysing the rules and statistics of the data sequences inside the events. This step is aimed for finding

6.1.   Rules (including probabilistic) of temporal sequences of particular meteorological parameters.

6.2.    Correlations and/or functional relationships between different meteorological parameters.

 

The above pp. 1 to 6 are preparatory steps, i.e. meteorological data analysis. Although the analytic part of the proposed method seems rather expanded, one may consider that it must be employed only once.

 

Synthetic sequence of operations

After the analysis of the meteorological data we propose the forcing data generation scheme in three basic steps, see also fig. 2.

 

7.      Generation of the artificial sequence of events. The event sequence is generated basing on the previously determined rules about (1) the sequence of events, i.e. probability distribution of transfer from one event type to another, and (2) the probabilities of the event duration. The average probabilities may be used to generate the typical sequence of events. The probability distributions may be modified (weighted) to obtain different (say, critical) artificial data sets (primary moderation). Thus, for critical data sets one may (1) increase the probabilities of transition to more relevant (in sense of system response, as aquired by process mapping) events; (2) increase the probabilities that these relevant events last longer.

8.      Generation of the artificial sequences of [primary] meteorological data inside the particular events. The data sequences are generated from the previously determined rules about the probability distributions of the particular quantitative parameters defining the meteorological data sequence inside the event. The data items forming the actual meteorological conditions are not independent; therefore the generation of the data sequence inside the event should be done for selected [primary] parameter: either most relevant parameter for further prognostic modelling, or parameter clearly determining other meteorological data. The probability distributions of the quantitative descriptors may be modified to increase/decrease the severeness and/or relevance of the particular event (secondary moderation).

9.      The building of the artificial sequence of the meteorological (or forcing, or model input) data by filling the sequences of p.8 into events of p.7.

 

Thus, sequence pp. 7 to 9 ensures the generation of nearly – natural forcing data set. The data set is typical if the average probabilities are used where required. The data set may be turned more/less critical by both/either (1) changing the occurrence/duration of events with different relevance, (2) changing the probabilistic parameters, determining the sequence of the data within particular event.

 

 

 

Application of proposed approach

 

Aim of study and available data

The proposed approach was applied for providing artificial data sets for longshore load transport modelling (Grzibovskis et. al., 1999). The time-series of the wind measurements (velocity and direction) was primarily required for numerical simulations. Far sea wave data was calculated by fetch model (Sennikovs et. al., 1998), whilst water level considered as less significant, and found from the wind parameters via correlation analysis. Seasonal, annual and long-term data series were required.

 

Continuous data series for six years (LHMA, 1992-97) of almost full set of meteorological observations at Riga station (air temperature, humidity, pressure, wind speed and direction, total and high cloudiness, precipitation), synoptic pressure charts (LHMA, 1985-97) for thirteen years, and wind data since 1971 were available for the analysis. Meteorological observations were recorded 8 times daily, whilst pressure charts were issued 5 times per week (on workdays).

 

Event selection, classification and expert recognition

The primary data required for morphodynamic calculations is wind data. However, we failed to identify sub-sets of wind data to qualify as events. The criterions applied for event selection were (1) an event must be simple but meaningful, (2) contents of the weather conditions must follow some general rules within an event, (3) duration of an event must be in synoptic time scale, (4) weather conditions during the actual event must be as independent as possible of preceding events, i.e. development of weather conditions should be self-containing within an event. All above criterions hold true for the time periods, associated with the passage of [particular periphery of] some baric formation (cyclon L, or anticyclon H) over the place of observation.

 

The classification of events was assumed according to the type of baric formation (cyclon L, anticyclon H, or neutral formations N) and the periphery passing over the observation place (8 peripheries for L and H). The recognition of the baric type was assumed possible from the air pressure, whilst periphery – from the wind time-series. Region of genesis (Icelandic, South European, local for L; Azoric, Greenlandic, East European, South European or local for H) was also considered as classification criteria. However, the recognition of the region of genesis is possible mainly via analysis of a sequence of pressure charts as well as from the seasonal temperature and humidity characteristics. The automated analysis of graphical information was considered too complicated, whilst temperature and humidity was unrelevant for the particular application.

 

The air pressure was assumed as the basic criteria for recognising of the baric formation. Formation was assumed as L, if Po £ 1010 hPa, and as H, if Po ³ 1015 hPa. Otherwise the conditions were assumed as neutral (except for short transitions between H and L). Other basic formations (local depressions, appendixes of H or L, transition regions) were not considered as independent.

 

The expert (“manual”) recognition of the events were performed for the whole 6 years time period, identifying also the region of genesis. 31,7% of total time may be associated with cyclons, 45,5% with anticyclons, whilst 22,8% with neutral conditions. Most (68%) of cyclons were Icelandic, generally moving from W to E. These cyclons mainly passed the Northern part of Baltic sea during the time period under consideration; thus, the location of observations was mainly influenced by their southern peripheries, where the warm sector of cyclon is rather typically residing. Thus, quite variable (although typical!) sequences of observations were characteristic for cyclons. The occurrence of the anticyclons of different origin was rather uniform (12 to 20% of total amount of H).

 

Automated recognition of events

The direct algorithm of the recognition of the baric formations was based on the same air pressure criteria as expert recognition. We assumed L for Po £ 1010 hPa, and H for Po ³ 1015 hPa. The algorithm was extended by

 

·        dividing transient periods between preceding and following events instead of assuming them neutral if transition time is less than 24 h;

·        eliminating weak cyclons (Po > 1008 hPa during whole lifetime of cyclon) and anticyclons (Po < 1018,5 hPa or Po < 1020 hPa and W £ 3 m/s during whole lifetime of anticyclon).

 

The periphery of the baric formation was detected according to the following algorithm:

 

·        the average wind direction for the beginning (12 h), end (12 h) and whole event was calculated;

·        the periphery of the cyclons was detected by subtracting 60°, whilst periphery of anticyclons by adding 150° from/to the average wind direction during the events;

·        the turning angle a of the wind was used for the detection of whether the center of baric formation passes the place of observation. We voluntary supposed that it is the case, if ½a½ > 135°. The comparison of the expert and automated event recognition is given in table 1.

 

Table 1 Comparison of expert and automated event recognition

 

H

L

N

Expert

Automatic

Expert

Automatic

Expert

Automatic

General fit, %

-

97,2

-

98,6

-

83,7

Average duration, h

98,0

131,0

81,5

79,6

59,0

71,9

Periphery accuracy 90 ° and better, %

-

91,7

-

95,6

-

-

 

Generally, most (97 to 99 % of H/L) of observations were found as belonging to the “right” baric formations. Disagreement between the expert and automatic event recognition is associated with weakly-defined formations. The comparison of average event duration indicated that direct algorithm fails to recognize sequential passing of several anticyclons, when air pressure remains high also during the transition period. The typical wind speeds during these anticyclons are rather low, therefore the wind data analysis does not improve situation; on the other hand the expected relevance of the anticyclonic events is low.

 

The “accuracy” of the determining of periphery is about 90° by expert evaluation, whilst 45° was used for the automated periphery recognition. One may consider the performance of direct algorithm as rather good. Note, that more distinct (short and strong) cyclonic events are processed better.

 

The attempts of applying AI approach for event recognition lead to less reasonable results, and they are omitted in this paper.

 

Morphodynamic relevance of events

The morphodynamic relevance of different events was assessed by mapping them to the integral longstore load transport through single cross shore profile at Latvian coast (orientation from SSW to NNE). The average load transport (per hour) in both direction during 17 types of events is summarised in table 2.

 

One may conclude that (1) the same events may be responsible for sediment movement in both directions due wind turns during the event, (2) anticyclons have lower morphodynamic relevance as cyclons, (3) only two to three peripheries of both cyclons and anticyclons (i. e. ~ 30% of events) may be of morphodynamical significance.


Table 2 Morphodynamic relevance of different events

Event type

QN, m3/h

Qs, m3/h

Q, m3/h

L-N

L-NE

L-E

L-SE

L-S

L-SW

L-W

L-NW

1

0

4

17

123

67

2

1

-4

-1

-2

-5

-47

-321

-88

-27

3

-1

2

12

76

-254

-86

-26

H-N

H-NE

H-E

H-SE

H-S

H-SW

H-W

H-NW

3

10

4

1

0

0

1

1

-10

-32

-34

-38

-8

-2

-6

-2

-7

-22

-30

-37

-8

-2

-5

-1

N

4

-9

-5

 

Generation of event sequences and wind data series

Two basic probability distribution groups (probability table of transition between different events, and probabilities of duration of different events) were calculated, and applied for obtaining random (but nearly-natural) event sequence. Both probability distributions are essentialy seasonal.

 

The principal modelling input data is wind speed and direction. Therefore the analysis of the internal rules of the wind data inside the events was performed. Principal simplification was applied assuming that wind speed and direction may be considered independently.

 

The ability of recognition of the event from changing wind direction (turning) was demonstrated before. Thus, we applied the general rule of changing wind direction clockwise or anticlockwise during the passage of baric formation. The approach of describing wind direction time development by three angles (a1 at the start moment, a2 at the middle, a3 at the end moment of event) was used. 17 ´ 3 probability distributions , were calculated for all event types ´´.

 

The following general rules of changing wind velocity during the event were applied:

1.      The wind speed in the center of H is lower as in periphery (see typical example in fig. 3).

2.      The wind speed in the center and peripheries of cyclons is lower as in some [circular] intermediate zone (see example in fig. 4).

3.      The wind speed is rather low during neutral events, whilst it is either falling or raising during transition between events.

 

First two rules allow to distinguish three principally different time-developments of wind speed within an event:

1.      Wind speed has one minimum during the passage of the anticyclon (as in fig. 3).

2.      Wind speed has two maximums and central minimum, if the central part of the cyclon passes close to the observation point (as in fig. 4).

3.      Wind speed has one maximum during the passage of the cyclon periphery.

 

Fig.3 Wind speed (m/s) during passage of (eastern periphery) of anticyclon, May, 1992; time in days

 

We applied the following scheme of the generation of the wind speed sequences. Analytic part (time and wind speed normalized to event duration and maximum wind speed):

1.      Approximation of the wind speed time development by 4th order polynomial expression for all recognized events by least square method.

2.      Finding 5 characteristic nodes (ti, Wi), i = 1, 2… 5 for each event: (1) wind speed at the beginning and end of events, (2) wind speeds and their occurrence times at wind speed extremes N £ 3, (3) 5 – N – 2 wind speed and time values in intermediate time moments.

3.      Calculation of 8 probability distributions for each event type: , i = 2, 3, 4; ,    i = 1, 2… 5 (´´ is event type).

4.      Calculation of the probability distributions for maximum dimensional wind speed  and the non-dimensional deviation of the maximum of approximation curve from the observed wind speed maximum .

 

Synthetic part, i.e. generation of wind speed time development during the particular event was performed by the following algorithm:

5.      Random determination of the event’s strength Wmax from the probability distribution .

6.      Random determination of the non-dimensional time moments t2, t3, t4 and non-dimensional wind speeds W1, W2, W3, W4, W5 from the probability distributions ,    i = 2, 3, 4 and , i = 1, 2… 5.

7.      Construction of polynomial interpolation curve through five characteristic points (ti, Wi).

8.      Random perturbation of the discrete analogue of the interpolation curve according to probability distribution .

9.      Transition to dimensional time and wind speed.

 

Fig.4 Wind speed (m/s) during passage of center and SW periphery of cyclon, Feb-1992; time in days

 

 

Above approach resulted in artificial whilst nearly natural time-series of wind data, further employed for the prognostic morphodynamic calculations.

 

 

Conclusions

The proposed approach to the knowledge of authors is rather innovative and may be widely employed for providing input data required for a range of different prognostic calculations.

 

At the same time (1) particular steps of the proposed method should be carefully tested and improved as well as (2) procedures for the quality assessment of the generated data sets should be developed and employed. Both lather issues regretfully fall outside the scope of this paper.

 

References

R.Gržibovskis, U.Bethers, J.Sennikovs 1999. Modelling littoral drift along Latvian and Germany coasts”. Proc. International Scientific Coll., Modelling of Material Processing”, Riga, pp. 180 - 185.

J. Senņikovs, R. Gržibovskis, U.Bethers, K.-P. Holz 1998. Multi-level approach to the estimation of the load transport near Ruegen. Proceedings ICHE'98, Cottbus. Abstract in “Advances in Hydro – Science and Engineering”, Vol. III, The University of Mississippi, (p.124). Full text on CD “CyberProceeding 3rd Int. Conf. on Hydroscience and Engineering”, ISBN 0 – 937099 – 09 – 0.

LHMA 1992-97. Tables TMS. Technical reports. Hydrometeocenter of the State Hydrometeorological Agency of Latvia. In Latvian.

LHMA 1985-97. Meteorological bulletins of the Latvian Hydrometeorological Agency. Hydrometeocenter of the State Hydrometeorological Agency of Latvia. In Latvian.