1. Introduction
We discuss two use cases that show how the GrammarViz GUI discovers time series patterns of variable length.
2. Datasets used
Two datasets are used in this demo:
2.1. Winding dataset
This dataset is a snapshot of data collected from an industrial winding process (the DaISy collection’s test setup of an industrial winding process); its column #2 corresponds to the traction reel angular speed:

2.2. QTDB 0606 ECG dataset
The database record can be downloaded from the PhysioNet QT Database and converted into text format with
rdsamp -r sele0606 -f 120.000 -l 60.000 -p -c | sed -n '701,3000p' >0606.csv
(assuming rdsamp from the WFDB toolkit is installed on your system). We use the second column of this file; a ready-made copy ships with GrammarViz as data/ecg0606_1.csv. This is our dataset overview:

We know that the third heartbeat of this dataset contains the true anomaly, as discussed in the HOT SAX paper by Eamonn Keogh, Jessica Lin, and Ada Fu — see the anomaly discovery tutorials for that side of the story. Here, we are after the recurring patterns.
3. Time series discretization and grammar induction
There are two ways to perform the discretization: globally, or with an overlapping sliding window (toggled by the “Slide the window” checkbox). We use sliding window-based discretization in this demo.
The discretization is configured by three numeric parameters: the sliding window size, the PAA size, and the alphabet size. These control the granularity of the approximation, which in turn drives the sensitivity and selectivity of the higher-level algorithms — for example, the ability to capture a local phenomenon. Note, however, that the grammar induction step effectively mitigates a suboptimal sliding window choice, because rules compose across windows. The grammar induction itself requires no parameters.
4. Variable-length recurrent pattern discovery
We use the winding dataset in this example. Click “Browse…”, select winding.csv, then “Load data” — the GUI plots the winding data. Now set the discretization parameters: window size 100, PAA size 4, alphabet size 3. Click “Discretize”. At this point the GUI displays the inferred grammar rules.
Click the “Mean length” column header to sort the rule table in ascending order; click it again for descending. Select the rule with the longest mean length — the GUI shows that this rule spans about 353 points and is observed twice:

Similar long rules are also observed twice each, at lengths of roughly 227 and 187 points.
Now click the “Frequency in R0” column header twice and choose the most frequent rule:

The subsequences corresponding to this rule range in length from 102 to 121 points, and a few of them overlap.
Note that these variable-length subsequences correspond to a single rule of the same grammar inferred by Sequitur from the discretized time series.
Similarly, if we load the qtdb0606 ECG dataset with the discretization parameters set to sliding window 100, PAA 8, and alphabet 4, the algorithm finds that the most frequently occurring rules are the normal heartbeats:

5. Discussion
Note that the recurrent pattern discovery technique shown above yields sets of frequent subsequences of variable length. Two factors make this possible: the numerosity reduction embedded in the discretization step, and the nature of GI algorithms, which build their rules from long-range correlations in the input.