Enter a job title or keyword

Splunk Observability Engineer

Diligent Tec Inc


Job Location:

Chicago, IL - USA

Monthly Salary: Not provided by the employer
Posted: 22 August 2026 (20 hours ago)
Application Deadline: 19 November 2026
Vacancies: 1 Vacancy

Job Summary

Location: Chicago IL onsite

To design implement and optimize a full-stack observability strategy using theSplunk Observability Cloud (formerly SignalFx)andSplunk Enterprise/Cloud. You will ensure that engineering teams have 360-degree visibility into system health moving the organization from reactive firefighting to proactive pattern-based incident prevention.

Key Responsibilities
Data Orchestration:Architect the ingestion of the Three Pillars (Metrics Logs Traces) usingOpenTelemetry (OTel)collectors.
Aggregation Strategy:Develop logic to aggregate high-cardinality data to reduce noise while maintaining signal for troubleshooting.
Analytical Modeling:UseSPL (Search Processing Language)andSignalFlowto perform pattern analysis detecting anomalies before they trigger traditional threshold alerts.
Visual Storytelling:Build executive and technical dashboards that correlate disparate data points (e.g. showing how a spike in 500-errors inLogsrelates to a specific span in aTrace).
Required Hands on Technical Skills
1. Telemetry & Data Specialization
Logs:Proficiency in Logging-in-Context. You must be able to link logs directly to trace IDs so developers can jump from a failing trace to the specific line of code in the logs.
Metrics:Expertise inSignalFlow(Splunks background streaming analytics language). You should know how to calculate percentiles ($P95 P99$) rates of change and historical averages.
Traces:Deep understanding ofDistributed Tracing. You must know how to instrument applications (Java Python Go) to capture spans and identify bottlenecks in microservices.
2. Pattern Analysis & Aggregation
Anomaly Detection:Ability to configureMetric FinderandMDetectorusing standard deviations or Mean Absolute Deviation to find outliers.
Data Scrubbing:Skills in usingSplunk Ingest ActionsorEdge Processorsto filter mask or aggregate dataat the edgeto save on license costs and improve search speed.
Pattern Discovery:Using Splunks machine learning commands () to group millions of log events into a few dozen patterns for faster root cause analysis.
3. Hands on - Dashboards & Visualization
High-Cardinality Handling:Designing dashboards that dont break when viewing thousands of containers.
Contextual Drill-downs:Building Glass Tables (in ITSI) or Unified Dashboards that allow a user to click a metric and immediately see the associated logs.
Frameworks:Familiarity with theDashboard Studioand JSON-based dashboard definitions for version control (GitOps).
Preferred Qualifications & Certifications
DevOps & IAC skills
Splunk Cloud Certified Metrics User:Focuses on the metrics and alerting side.
Splunk Core Certified Power User:Essential for mastering complex SPL for log analysis.
OpenTelemetry Expert:Knowledge of theOTel Collectorconfiguration (receiversprocessorsexporters) is currently the most in-demand skill for this role.
Job Specification: Splunk Observability Engineer
Chicago IL onsite

To design implement and optimize a full-stack observability strategy using theSplunk Observability Cloud (formerly SignalFx)andSplunk Enterprise/Cloud. You will ensure that engineering teams have 360-degree visibility into system health moving the organization from reactive firefighting to proactive pattern-based incident prevention.

Key Responsibilities
Data Orchestration:Architect the ingestion of the Three Pillars (Metrics Logs Traces) usingOpenTelemetry (OTel)collectors.
Aggregation Strategy:Develop logic to aggregate high-cardinality data to reduce noise while maintaining signal for troubleshooting.
Analytical Modeling:UseSPL (Search Processing Language)andSignalFlowto perform pattern analysis detecting anomalies before they trigger traditional threshold alerts.
Visual Storytelling:Build executive and technical dashboards that correlate disparate data points (e.g. showing how a spike in 500-errors inLogsrelates to a specific span in aTrace).
Required Hands on Technical Skills
1. Telemetry & Data Specialization
Logs:Proficiency in Logging-in-Context. You must be able to link logs directly to trace IDs so developers can jump from a failing trace to the specific line of code in the logs.
Metrics:Expertise inSignalFlow(Splunks background streaming analytics language). You should know how to calculate percentiles ($P95 P99$) rates of change and historical averages.
Traces:Deep understanding ofDistributed Tracing. You must know how to instrument applications (Java Python Go) to capture spans and identify bottlenecks in microservices.
2. Pattern Analysis & Aggregation
Anomaly Detection:Ability to configureMetric FinderandMDetectorusing standard deviations or Mean Absolute Deviation to find outliers.
Data Scrubbing:Skills in usingSplunk Ingest ActionsorEdge Processorsto filter mask or aggregate dataat the edgeto save on license costs and improve search speed.
Pattern Discovery:Using Splunks machine learning commands () to group millions of log events into a few dozen patterns for faster root cause analysis.
3. Hands on - Dashboards & Visualization
High-Cardinality Handling:Designing dashboards that dont break when viewing thousands of containers.
Contextual Drill-downs:Building Glass Tables (in ITSI) or Unified Dashboards that allow a user to click a metric and immediately see the associated logs.
Frameworks:Familiarity with theDashboard Studioand JSON-based dashboard definitions for version control (GitOps).
Preferred Qualifications & Certifications
DevOps & IAC skills
Splunk Cloud Certified Metrics User:Focuses on the metrics and alerting side.
Splunk Core Certified Power User:Essential for mastering complex SPL for log analysis.
OpenTelemetry Expert:Knowledge of theOTel Collectorconfiguration (receiversprocessorsexporters) is currently the most in-demand skill for this role.

Required Skills:

SPLUNKPYTHONJAVAVISUALDATA SCRUBBINGMACHINE LEARNINGJSONDESIGNINGCAUSE ANALYSISVERSION CONTROL