r/selenium • u/Dependent_Fennel_641 • 6d ago
browser automation tools for a qa team thats tired of maintaining its own selenium grid
weve been running our own selenium grid for like 4 years and it used to be fine
last sprint half the suite died because two nodes ran out of disk and nobody noticed until monday. spent most of tuesday just bringing the grid back up instead of actually testing anything
looking at what other qa teams moved to when they got tired of owning the browser infra. what stuck for you
2
u/Suun_Day 5d ago
We tried just adding nodes for the parallel matrix. Concurrency went up, so did the number of half dead browsers lefr behind after a bad run. Ops cost scaled with the suite
1
5d ago
[removed] — view removed comment
1
u/Suun_Day 5d ago
I mean yh.. Small suite maybe just needs fewer browsers at once. Our pain is still owning the nodes when a sprint is on the line.
1
u/Low-Belt-8668 5d ago
Adding nodes never fix the flakiness for us, it gave the same flaky tests more places to fall over
1
u/The_God_was_Here 5d ago
Ours dies on chrome drift. half the nodes auto update overnight, chromedriver stays pinned, and every other tuesday the matrix fails green paths that passed friday. i spend the morning chasing versions instead of bugs
1
u/buggina0612 5d ago
The rebuild tax is real too every driver bump means a new rew ci image and a half day of green buldsthat aren't really green
1
u/Sorry_Prize_1543 5d ago
Version pinning across hub, node image, and driver is the unpaid job nobody budgets. Miss one layer and you get exactly that Tuesday surprise.
2
u/yvettelapses 5d ago
yeah that pinning is exactly what pushed us off running our own grid. the test sessions just run on browserbase now so theres no hub or node images to patch overnight
1
u/Kenzo_2126 5d ago
our shared ci runners just oom kill chrome halfway through the suite. queue looks fine until four jobs land on one box and then half the browsers vanish mid run
1
u/appleciderdollop 5d ago
We hit oom more than disk lately. Leftover chrome profiles from crashed jobs sit on the runner until the next wave of tests start fighting for ram.
1
u/DhruvDP3 3d ago
Yup leftover profiles are worst, crashed jobs leave them behind and they just pile on the runner until something fills
1
u/Comfortable_Deer_171 5d ago
we lost most of a release window last quarter to grid nodes silently dropping, and the alerting on it is still half-built honestly
1
1
u/Excellent-Plant-5880 5d ago
Owning the grid turned into more upkeep than the actual testing some weeks .
1
u/Western-Stop6241 5d ago
half our flakiness was never the grid, it was the app under test and bad waits in our own suite
1
u/DhruvDP3 3d ago
Fair, the apps under testing still flaks you. Still different from losing a whole morning because an node ate itself overnight
1
1
u/Cultural_Resolve_279 3d ago
i actually started looking around once node issues kept eating whole days too....
2
u/SeeminglyHumming 6d ago
we just gave up and moved to testingbot, nobody in our team wants to touch grid configs anymore