Batch indexing FAQs
What is batch indexing?
Batch indexing refers to the background process which builds up the index data (historical tag data) for a specified list of tags.
Why do batch indexing?
TrendMiner syncs a list of available tags from a data source. But not all these tags are also indexed. Only tags which are used at least once are being indexed in TrendMiner.
When the tag is used for the first time TrendMiner needs to pull in the full historical data for the tag. Depending on the data source performance, the general load on the system and the index horizon this can take up to several minutes per tag (e.g. 1 second to fetch 1 month of data result into 1 minute per 5 years of indexing). As long as the index data is not available users cannot run their analytics on these tags.
To make sure that newly added tags are immediately available for analysis (no waiting time) an admin can trigger batch indexing.
How to trigger batch indexing?
TrendMiner provides a support script which can be run to batch index a list of tags. More information can be found here: Batch Indexing
How long does it take to run batch indexing?
Batch indexing is basically an automation to simulate tags being used. In the background the tag indexing performance is the same as if the tag would be used by an end user and therefore the indexing speed can be expected to be the same.
The speed of historical indexing is highly related to the data source performance and the load on the TrendMiner installation (mainly the number of already indexed tags and number of active monitors). On a healthy and performant system 1 tag can finish indexing for 5 years of historical data in about 1 minute.
On a heavily loaded or congested system you can expect tag indexing to progress with a speed of 1 month per hour (for a 5 year index horizon that's 1.67% progress per hour). The batch indexing will not progress faster to reduce the impact of this background process on the users and active monitors.
Side note: for historical indexing the 3 most recent months are processed with high priority. Older data is indexed with a lower priority so it can be expected that the index completion rate slows down shortly after a new tag is being indexed.
Can I speed up the batch indexing?
Batch indexing is meant to run as a low priority background process with minimal impact on the users and active monitors.
In case timely finishing of the batch indexing is required it can be considered to temporarily reduce the load on the TrendMiner system and the data sources to reserve maximal capacity for the batch indexing. This could be achieved for example by temporarily disabling all monitors and/or forward indexing. Since this is not considered standard practice some preparation and manual actions are required. Contact TrendMiner support to discuss the options.
Triggering a large amount of parallel indexing tasks (100+) at the same time can overload the system and lead to performance complaints from users and delayed monitor results. On top of that some data sources will show degraded performance and as an end result to total time to finish the batch indexing will increase compared to a lower amount of parallel indexing tasks.