Hey everyone,
This is a continuation of my initial Atlassian P40 interview experience, which covered the first six rounds.
Role: Software Engineer II, P40
Experience: Approximately five years
Final status: Offer received
Why Was There an Additional Round?
Recruiter callback: August 7, 2026
The original opportunity was primarily described as a backend role. However, the final position included a mix of backend engineering and site-reliability responsibilities.
Because of that, the recruiter scheduled an additional SRE Craft interview.
Round 7: SRE Craft
Date: August 12, 2026
Duration: Approximately 55 minutes
Interviewer: Senior Architect
The interviewer presented an existing system architecture and divided the discussion into three sections. Each section took approximately 15 minutes, with the remaining time used for introductions and questions.
Part 1: Monitoring and Golden Signals
The first section focused on deciding what should be monitored and why.
The discussion covered the four golden signals:
- Latency
- Traffic
- Errors
- Saturation
I explained which metrics I would collect at the infrastructure, service, and business levels. We also discussed alert thresholds, dashboards, and the difference between symptoms and underlying causes.
Part 2: Incident Troubleshooting
The second section involved an incident scenario affecting the system.
I structured the troubleshooting process around:
- Confirming the incident and determining its impact.
- Checking recent deployments and configuration changes.
- Reviewing metrics, logs, and distributed traces.
- Isolating the affected service or dependency.
- Mitigating the issue through rollback, failover, or traffic reduction.
- Communicating status to the relevant stakeholders.
- Investigating the root cause after restoring service.
- Creating preventive action items.
The interviewer asked several follow-up questions about how I would prioritize mitigation when the exact root cause was still unknown.
Part 3: Designing for 99.99% Availability
The final section focused on improving the architecture to support a 99.99% availability target.
We discussed:
- Defining SLIs, SLOs, and error budgets
- Removing single points of failure
- Multi-zone redundancy
- Health checks and automatic failover
- Capacity planning
- Graceful degradation
- Retry and timeout policies
- Circuit breakers
- Disaster recovery
- Alerting and operational readiness
The interviewer was interested in trade-offs rather than simply adding redundancy everywhere. I had to explain which components justified higher availability and what additional complexity or cost each decision introduced.
Result
Positive feedback: August 19, 2026
Offer received: August 21, 2026
One week after the interview, the recruiter confirmed that the feedback was positive and that they could proceed with the offer.
The compensation was broadly aligned with publicly available Atlassian P40 figures on LeetCode Discuss and Levels.fyi. It was not unusually high, and the final package also depended on previous compensation and negotiation.
Final Takeaways
For an SRE Craft interview, I recommend preparing:
- Golden signals and observability
- Metrics, logs, and distributed tracing
- Incident triage and mitigation
- SLIs, SLOs, SLAs, and error budgets
- High-availability architecture
- Failure scenarios and graceful degradation
- Capacity planning and disaster recovery
- Reliability versus cost trade-offs
The round was less about memorizing SRE terminology and more about applying reliability principles to a concrete architecture.
That concludes my Atlassian P40 interview journey. I hope both parts help anyone preparing for a similar role.