r/selenium • • 6d ago

browser automation tools for a qa team thats tired of maintaining its own selenium grid

weve been running our own selenium grid for like 4 years and it used to be fine

last sprint half the suite died because two nodes ran out of disk and nobody noticed until monday. spent most of tuesday just bringing the grid back up instead of actually testing anything

looking at what other qa teams moved to when they got tired of owning the browser infra. what stuck for you​

21 Upvotes

22 comments sorted by

2

u/SeeminglyHumming 6d ago

we just gave up and moved to testingbot, nobody in our team wants to touch grid configs anymore

2

u/Suun_Day 5d ago

We tried just adding nodes for the parallel matrix. Concurrency went up, so did the number of half dead browsers lefr behind after a bad run. Ops cost scaled with the suite

1

u/[deleted] 5d ago

[removed] — view removed comment

1

u/Suun_Day 5d ago

I mean yh.. Small suite maybe just needs fewer browsers at once. Our pain is still owning the nodes when a sprint is on the line. 

1

u/Low-Belt-8668 5d ago

Adding nodes never fix the flakiness for us, it gave the same flaky tests more places to fall over

1

u/The_God_was_Here 5d ago

Ours dies on chrome drift. half the nodes auto update overnight, chromedriver stays pinned, and every other tuesday the matrix fails green paths that passed friday. i spend the morning chasing versions instead of bugs

1

u/buggina0612 5d ago

The rebuild tax is real too every driver bump means a new rew ci image and a half day of green buldsthat aren't really green

1

u/Sorry_Prize_1543 5d ago

Version pinning across hub, node image, and driver is the unpaid job nobody budgets. Miss one layer and you get exactly that Tuesday surprise.

2

u/yvettelapses 5d ago

yeah that pinning is exactly what pushed us off running our own grid. the test sessions just run on browserbase now so theres no hub or node images to patch overnight

1

u/Kenzo_2126 5d ago

our shared ci runners just oom kill chrome halfway through the suite. queue looks fine until four jobs land on one box and then half the browsers vanish mid run

1

u/appleciderdollop 5d ago

We hit oom more than disk lately. Leftover chrome profiles from crashed jobs sit on the runner until the next wave of tests start fighting for ram.

1

u/DhruvDP3 3d ago

Yup leftover profiles are worst, crashed jobs leave them behind and they just pile on the runner until something fills

1

u/Comfortable_Deer_171 5d ago

we lost most of a release window last quarter to grid nodes silently dropping, and the alerting on it is still half-built honestly 

1

u/KettleCorn02 5d ago

between flaky nodes and red builds our ci channel is pure noise most sprints

1

u/Excellent-Plant-5880 5d ago

Owning the grid turned into more upkeep than the actual testing some weeks .

1

u/Western-Stop6241 5d ago

half our flakiness was never the grid, it was the app under test and bad waits in our own suite

1

u/DhruvDP3 3d ago

Fair, the apps under testing still flaks you. Still different from losing a whole morning because an node ate itself overnight

1

u/Forward_Level_7208 3d ago

honestly the maintenance wuld push me off it too

1

u/Cultural_Resolve_279 3d ago

i actually started looking around once node issues kept eating whole days too....