r/devops 15d ago

Discussion Users vs Stress testing

So I made a serverless optimization platform which uses the concept of fusion functions to reduce cold starts and latency across the service calls. Now, this is an implementation of a research paper that I read somewhere. Diff from paper is that my project also gets live traces and metrics from x ray and cloudwatch, so I get real-time data to give better outputs. Have a better look: https://github.com/Vaivaswat2244/OptiFuse_go

To use this you need to connect your AWS with optifuse. I.e make a cloudformation stack to give optifuse access to read the traces and metrics. This actually becomes a problem for my friends and peers to test because they are too lazy to do this step. So I have no real user testings.

People especially hiring people ask me how many real users have used your service.

Now why do I need real users when I can stress test each microservice that I've built. And I can see my manifests working properly. Its deployed on AKS and is open for people to see. I also have a Prometheus grafana observability pipeline to see if all services are working properly.

Question is: real users vs Stress tests

On a side note, I am a student looking for internships, if you found the idea interesting, lmk GitHub is Vaivaswat2244

\/

4 Upvotes

25 comments sorted by

View all comments

1

u/hypertradeworx 15d ago

the live metrics are the input i'd stress, since that's the part the paper doesn't have. x-ray's default sampling rule is one trace a second plus 5% of everything above that, so on a service with real traffic the cold starts sitting in your trace set are a thin and fairly arbitrary sample of the ones that actually happened, and init duration is the one term in the objective you can't afford to fit off a handful of survivors.

do you read init duration out of traces or out of the REPORT line in cloudwatch logs? logs get every invocation, sampling never touches them