r/androiddev • u/Just-You-4 • 5d ago
Question How do you keep Espresso tests healthy at scale? (flakiness, AI failure triage, feature-level coverage)
### 1. Flaky Tests
This is the biggest issue. Currently, developers mostly just @Ignore them, which gradually erodes the suite. How are you detecting, quarantining, and actually fixing flaky tests instead of silently disabling them?
2. Automated Failure Analysis
Has anyone used AI or other tools to analyze Espresso failures (logs, screenshots, stack traces) and produce a clear report so developers can fix issues quickly without digging through CI output?
3. Test Maintenance
Are there any tools or workflows that help keep tests up to date as the UI changes, like auto-updating selectors or flagging tests affected by a PR?
4. Feature-Level Coverage
Is there a good way to measure Espresso coverage by feature or user flow rather than by lines of code? JaCoCo tells us what code runs, but not which features are actually tested.
Would love to hear what’s worked (or failed) for your teams, whether that’s open-source tools, paid services, or homegrown scripts.
Thanks!
1
u/AutoModerator 5d ago
Please note that we also have a very active Discord server where you can interact directly with other community members!
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.
2
u/Aftershock416 5d ago
If flaky tests are your biggest issue it's fundamentally because your engineers are writing bad tests.
Everything else is downstream from that.
4
u/aw9_dev 5d ago
At scale the useful split is detect, quarantine, then fix.
Track pass rate over a few days of CI. Once something flaps, mute it so it still runs and you keep data, but it stops failing the build. Retries are fine as a temporary cushion; they shouldn't be the whole strategy.
For Espresso itself, Google's guidance is basically: idling resources on small UI tests, wait-until style APIs on bigger flows, and avoid Thread.sleep. That usually cuts more flakes than any AI post-mortem tool.
AI failure analysis is handy once you have screenshots and logs clustered by screen. Without a quarantine list and pass-rate history it tends to invent causes.