Instrument web services with counters, gauges, and latency histograms.
Project-based Internship Programme
Master Observability and SLOs in a Site Reliability Engineering Internship Project
This site reliability engineering online internship with certificate is a fee-based, project-based internship programme. This site reliability engineering online internship with certificate is a fee-based, project-based internship programme that focuses on production reliability engineering, service level metrics, and incident management. You instrument a Kubernetes service with Prometheus metrics, define user-centric SLOs and error budgets, configure multi-window burn rate alerts, execute controlled failure simulations, and author an operational blameless postmortem.
Decision note 01
Who this project fits and who it does not
A useful fit if…
- You want to understand how top tech companies maintain 99.9% uptime across complex distributed architectures.
- You value data-driven reliability: replacing arbitrary alert thresholds with mathematical error budget burn rates.
- You want to build an SRE portfolio project showcasing Prometheus instrumentation, Grafana dashboards, chaos testing, and postmortem authoring.
Choose another route if…
- You only want to write frontend user interfaces; our Frontend Development programme focuses directly on UI components.
- You want CI/CD deployment pipelines without production observability; our DevOps programme focuses on continuous delivery.
- You believe system reliability means zero failures ever occur; SRE accepts failure and manages it with error budgets.
Official assigned project
Kubernetes SLO/Observability Project
In modern distributed cloud platforms, systems will fail. Hardware breaks, network partitions occur, and software deployments introduce bugs. Site Reliability Engineering (SRE) is the discipline of treating operations as a software problem. In this SRE project, you take ownership of an application service running in Kubernetes. You instrument the code with Prometheus metrics, define user-centric Service Level Objectives (SLOs), quantify your service's error budget, build real-time Grafana dashboards, configure multi-window burn rate alerts, and conduct a controlled failure drill followed by a blameless incident postmortem.
Kubernetes SLO/Observability Project - Instrument a Kubernetes service, define SLOs, and build a tested observability and incident workflow.
Catalogue deliverables
- Instrumented Kubernetes application service exposing Prometheus RED metrics (Rate, Errors, Duration)
- Formal Service Level Objective (SLO) specification document defining SLIs, targets, and error budget policies
- Prometheus Alertmanager rules configured for multi-window burn rate detection rather than transient CPU spikes
- Comprehensive Grafana dashboard JSON layout tracking service availability, latency percentiles, and error budgets
- Chaos testing report and blameless postmortem incident review documenting controlled failures and recovery actions
How the project works
From service instrumentation to error budget burn rate alerting and blameless postmortems
Reliability engineering begins by measuring what matters to actual users. Rather than monitoring low-level machine metrics like host CPU percentages (which rarely correlate directly with user satisfaction), you instrument your application using the RED method: Request Rate, Error Count, and Request Duration (latency). You expose these metrics via a clean `/metrics` HTTP endpoint formatted for Prometheus scraping.
Service Level Objectives (SLOs) turn qualitative reliability goals into quantitative engineering contracts. You establish a Service Level Indicator (SLI) measuring the percentage of successful HTTP requests completed in under 300 milliseconds. You set a realistic SLO target (e.g. 99.5% over a rolling 30-day window), which mathematically establishes an 'error budget', the exact amount of unreliability your service is permitted to experience before feature releases are halted to prioritize stability.
Alerting must be actionable rather than noisy. Alert fatigue causes on-call engineers to ignore critical warnings. You replace naive instant-threshold alerts with Google SRE multi-window burn rate alerts. Your Alertmanager configuration alerts only when the service is consuming its 30-day error budget at a rate that threatens to exhaust the entire budget within hours or days, filtering out transient, self-resolving network hiccups.
Failure testing and blameless postmortems close the reliability loop. You run a controlled failure exercise using a load testing tool (like k6), injecting artificial latency or killing cluster pods. You observe your alerts fire in Prometheus, monitor how Kubernetes restarts unhealthy containers automatically, and observe the alert resolve. Finally, you write a structured, blameless postmortem detailing the incident timeline, root cause, user impact, and permanent preventive engineering items.
Submission evidence
What makes this work reviewable
- Kubernetes deployment and service YAML manifests featuring Prometheus scrape annotations.
- Application source code demonstrating Prometheus client metric instrumentation (RED method).
- Formal SLO and error budget calculation documentation with mathematical justification.
- Prometheus Alertmanager rules configuration file defining multi-window burn rate alerting.
- Chaos injection execution log, Grafana dashboard export, and completed blameless postmortem incident report.
Your build path
Move from question to reviewable evidence
Instrument a containerized service with Prometheus client metrics, capturing HTTP traffic rates, errors, and latencies. Add Prometheus RED metrics (Rate, Errors, Duration) to a containerized microservice running in Kubernetes.
Formulate formal Service Level Indicators (SLIs) and establish a defensible 99.5% Service Level Objective (SLO). Formulate user-centric SLIs, establish a 99.5% availability objective, and calculate the 30-day error budget.
Build a Grafana observability dashboard displaying real-time error budget burn rates and latency histograms. Build a Grafana dashboard visualizing request rates, p95/p99 latency histograms, and remaining error budget.
Configure Prometheus alerting rules that trigger pages based on rapid error budget depletion. Author Prometheus Alertmanager rules detecting rapid error budget burn rates to eliminate alert noise.
Execute a controlled chaos injection exercise (e.g. synthetic load and pod failure), test automated recovery, and author a postmortem. Inject controlled cluster failures, observe alert firing and automated pod recovery, and author a blameless postmortem.
Private self-check
Is this project a reasonable learning fit?
Your answers remain in this browser tab and are not stored or sent.
Use these prompts for reflection; they are not an eligibility test.
Skills notebook
Build capability in a realistic order
These are general domain-learning suggestions, not confirmed HireeBridge tool requirements.
Observability & Instrumentation
Author time-series queries computing rates, quantiles, and rolling availability.
Design clean operational dashboards displaying latency heatmaps and budget burns.
Service Level Engineering
Define quantitative user-centric reliability targets based on business criticality.
Calculate allowable downtime budgets and establish release-freeze policies.
Implement multi-window alerting to catch rapid budget depletion while ignoring noise.
Resilience & Incident Management
Inject synthetic latency, resource starvation, and pod terminations safely.
Configure Kubernetes restart policies and health probes to restore crashed services.
Document incident timelines, root causes, detection gaps, and preventive action items.
Review before submitting
Common Site Reliability Engineering project mistakes
- 01
Alerting on transient CPU or memory spikes instead of user impact
CPU can spike to 90% during healthy batch jobs; alert only when user-facing requests are failing or exceeding latency SLOs.
- 02
Using average latency instead of 95th/99th percentiles
Average latency conceals severe lag experienced by unlucky users; always measure p95 and p99 percentiles.
- 03
Writing a postmortem that blames human error
Human error is a symptom, not the root cause; identify why the system allowed the mistake and build automated safeguards.
- 04
Setting unrealistic 100% availability SLOs
100% reliability is impossible and prevents product innovation; choose realistic targets like 99.5% with managed error budgets.
- 05
Failing to verify that alerting rules actually fire under stress
Always execute a controlled failure injection test to confirm Alertmanager fires and resolves as designed.
What reviewers check
Completeness against the assigned brief and deliverables; functional correctness; domain-relevant logic, data, metrics or implementation; required edge cases and failure handling; reproducible setup and submission evidence; and clear documentation of the completed work.
Reviewer
GreyRocks team
Verify metrics, alert firing/resolution, SLO calculations, and health behavior during a controlled failure.
Evidence language
Draft an honest CV bullet
Keep placeholders until you can replace them with evidence from your own project.
- Instrumented a containerized Kubernetes microservice with Prometheus RED metrics, tracking p95/p99 latency percentiles.
- Engineered a user-centric 99.5% SLO and error budget framework, reducing un-actionable on-call alert noise by [percentage].
- Constructed multi-window burn rate alert rules in Prometheus Alertmanager to detect rapid error budget depletion.
- Conducted chaos engineering experiments injecting synthetic latency, validating automated Kubernetes pod recovery.
- Authored a comprehensive blameless postmortem report with root cause analysis and preventative architectural remediations.
Project readiness
Prepare a strong project submission
Certificate and verification
Completion comes before the credential
GreyRocks serves as the independent technical evaluation and credential verification entity for HireeBridge programmes. Programme enrolment grants access to the project specification, Kubernetes starter manifests, and evaluation criteria; it does not automatically award a completion certificate upon payment alone. To receive certification, you submit your instrumented service code, Grafana dashboard JSON, Alertmanager rules, and completed blameless postmortem report. An SRE evaluator reviews your SLI definitions, PromQL formulas, burn rate logic, and postmortem thoroughness. Approved projects receive an official credential featuring a unique credential ID and QR verification link on GreyRocks.
- Complete
- Submit
- Review
- Approval
- Credential ID and QR
Read the certificate process · Verify a credential on GreyRocks
Duration: 1 Month / 4 Weeks.
Plan inclusions: Each domain maps to an assigned project and task specification. Reference repositories and comprehensive materials depend on the selected plan; certificates follow task submission and explicit reviewer approval.
Questions from students
Site Reliability Engineering internship FAQ
What is the difference between Site Reliability Engineering (SRE) and DevOps?
DevOps focuses on continuous integration, build automation, containerization, and delivery pipelines. SRE focuses on production reliability, observability (Prometheus/Grafana), defining user-centric SLOs and error budgets, managing incidents, and conducting blameless postmortems.
Do I need an expensive multi-node cloud cluster to complete this project?
No. You can run the entire Kubernetes cluster locally using free lightweight tools like Minikube or Kind (Kubernetes in Docker), running Prometheus and Grafana seamlessly on your development machine.
What is an error budget and why is it important?
An error budget is the allowable amount of downtime or failed requests permitted by your SLO (e.g. 0.5% for a 99.5% SLO). It provides a data-driven balance between shipping new features quickly and maintaining system stability.
What are the RED metrics?
The RED method focuses on three core user-centric indicators: Rate (number of requests per second), Errors (number of failing requests), and Duration (the amount of time requests take to complete).
How do I simulate a failure in my project?
You use a load testing tool (such as k6 or Apache Bench) to generate synthetic traffic, while introducing artificial latency or resource constraints into your application code, causing error budgets to burn and triggering alerts.
What is a blameless postmortem?
A blameless postmortem is an engineering retrospective that analyzes an outage without blaming individuals. It examines the timeline, technical root cause, and systemic improvements needed to prevent the failure from reoccurring.
How long does the programme take to finish?
The curriculum is designed for 4 weeks of structured, self-paced progress: week 1 covers Prometheus instrumentation, week 2 covers SLOs and Grafana, week 3 covers burn rate alerting, and week 4 executes failure testing and the postmortem.
How can employers verify my SRE certificate?
Each certificate features an official GreyRocks credential ID and a scannable QR verification code that displays your verified project scope and completion record on the online verification portal.
Next step
Choose your plan and start building.
Review plan details, included resources and the assigned project scope before you begin.