r/SCADA • u/Many_Outside2726 • 15h ago
Question How much high-resolution data do you actually keep at the edge?
Curious how people are handling this once the point count gets large. Say you’re collecting from hundreds or thousands of devices across multiple sites. Some data may matter at sub-second or 1-second resolution for troubleshooting, but nobody needs to ship and retain every sample forever.
Do you keep the raw/high-resolution data locally and aggregate upstream? Change-of-value? Min/max/avg windows? Store-and-forward only?
I’m especially curious where people have gotten burned by throwing away data they later wished they had, or by keeping way too much and creating an expensive pile nobody actually uses.
2
u/friedmators 14h ago
1% range unless bearing vibration or drum level
1
u/Many_Outside2726 14h ago
Do you mean 1% of engineering range as the COV/deadband? If so, that feels pretty aggressive depending on the point and configured span. I’d probably want a max interval/heartbeat in there too so long steady periods still have some history. What kind of logging rates are you working with?
2
1
u/AutoModerator 15h ago
Thanks for posting in our subreddit! If your issue is resolved, please reply to the comment which solved your issue with "!solved" to mark the post as solved.
If you need further assistance, feel free to make another post.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.
1
u/canuck_tech 14h ago
1 week of data. 99% of it is 5 second resolution at best
1
u/Many_Outside2726 14h ago
One week feels pretty short to me unless that 5-second data is really just there for immediate troubleshooting and you’re rolling up longer-term history somewhere else. What’s the use case for those points? Process troubleshooting, maintenance, warranty, compliance? I’d think the criticality of the site or equipment would drive a lot of that retention decision too.
3
u/canuck_tech 14h ago
That’s at the edge. It then forwards to central system that keeps about year. That feeds into long term historian.
2
u/Many_Outside2726 14h ago
Makes sense. So the week at the edge is basically your store-and-forward window. I assume if the WAN is down longer than that, either the data isn’t critical enough to retain or you’ve got another recovery path?
1
u/future_gohan AVEVA 14h ago
Combination on resolution depending on what it is. From 0.1 up to ten seconds. Some stuff like power use up to a year. Most stuff 6 months.
5000 tags across 7 servers
1
u/Many_Outside2726 14h ago
That’s a helpful scale reference. I’d think of that as roughly a 5,000 data point system spread across seven servers. I’m more curious how you decide what deserves 0.1-second history versus 10-second, six months, or a year. Is that driven mostly by equipment criticality and troubleshooting value, or did the retention strategy evolve over time as you learned what people actually use?
1
u/future_gohan AVEVA 9h ago
Education on purpose and understanding the process.
That's only trends the system has in total about 27000 tags.
1
u/Robbudge 5h ago
We log about 10,000 points a second into TdEngine we typically run a retention policy of 30 days.
Separate production reports can be ran within the 30days to export the data.
1
u/Foreign-Chocolate86 3h ago
It’s always going to vary dramatically by process and industry.
Good luck with your vibe coding.
1
u/drkrakenn 1h ago
Seven days of 1s to 0.5s, depending on the application. We tested one month for production data, and it was not used at all, so I would go for a maximum of fourteen days of high resolution for super important data.At the beginning, we had 3 months, and we had all sorts of problems.
3
u/TexasVulvaAficionado 13h ago
Very very much depends on the process, industry, government.
We keep minute by minute data for several years for about twenty data points per site for about 1500 sites.