Cisco Logo

Cisco

Software Engineering Tech Lead (SRE + AI)

Posted 7 Days Ago
Be an Early Applicant
In-Office
London, Greater London, England
Senior level
In-Office
London, Greater London, England
Senior level
Lead the architecture and development of an AI-powered Production Intelligence platform combining SRE, observability, agentic AI, and automated remediation. Build AI agents, MCP integrations, telemetry pipelines, anomaly detection, evaluation frameworks, and safe human-in-the-loop workflows. Define reliability practices including SLIs, SLOs, error budgets, incident response, and post-incident reviews. Mentor engineers, establish technical standards, and coordinate delivery across global application and infrastructure teams.
The summary above was generated by AI
Meet the Team

The Collaboration Technology Group is redefining the future of teamwork, building services that connect people effortlessly across devices, locations and time zones.

Our team builds, runs and continuously improves the platform services behind Cisco’s collaboration products, operating at global scale across numerous datacentres. We’re a passionate, collaborative team focused on reliability, innovation and engineering excellence.

Your impact

As a Technical Leader, you will drive the architectural vision and implementation of our next-generation AI-powered Production Intelligence platform. You will combine Site Reliability Engineering practices with modern agentic AI (reusable Skills, Model Context Protocol (MCP), and LLM tooling) to transform how engineering and leadership teams monitor, diagnose, and auto-remediate global SaaS infrastructure.

What you'll do
  • Technical Leadership & Architecture: Define the technical roadmap and architecture for AI-assisted observability, automated incident response, and self-healing cloud infrastructure.
  • Agentic Workflows & Tooling: Design and build production-grade AI agents, MCP tool integrations, and deterministic evaluation pipelines for automated operational decision support.
  • Telemetry & Insights: Architect ingestion and correlation pipelines across distributed logs, metrics, OpenTelemetry traces, change events, and runbooks to accelerate Mean Time to Detection (MTTD) and Resolution (MTTR).
  • Safe Production Automation: Develop proactive anomaly detection and Human-in-the-Loop (HITL) remediation workflows with rigorous safety, security, and quality guardrails.
  • Reliability & Scalability Engineering: Partner with application and infrastructure teams to define SLIs/SLOs, handle error budgets, and lead deep-dive post-incident reviews (PIRs).
  • Mentorship & Collaboration: Mentor senior and mid-level engineers, establish engineering best practices, and drive alignment across global development and operations teams.
  • You’ll manage priorities and deadlines, communicate progress clearly and work across teams to turn production needs into reliable software and AI-assisted capabilities.
Minimum qualifications
  • Bachelor’s degree + 8 years of related experience, Master’s + 6 years, or PhD + 3 years in Computer Science, Software Engineering, or a related technical field.
  • Proven record as a Technical Lead or Lead SRE/Software Engineer delivering distributed, high-availability SaaS platforms at scale.
  • Strong proficiency in Python, Go, Java, or C++ with experience designing microservices, APIs, and production automation.
  • Deep experience with Kubernetes, Docker, and container orchestration in large-scale multi-cluster environments.
  • Proven background in SRE practices: SLI/SLO design, observability platforms (metrics/logs/traces), incident management, and automated RCA.
Preferred Qualifications
  • AI & Agentic Systems: Hands-on experience building LLM pipelines, AI Agents, Model Context Protocol (MCP) servers/clients, RAG architectures, and evaluation frameworks.
  • Observability & Telemetry: Experience with OpenTelemetry (OTel), Prometheus, Grafana, Splunk, ThousandEyes, or distributed tracing systems.
  • Cloud & Infrastructure: Expertise in public cloud providers (AWS, GCP, Azure), Terraform/IaC, and GitOps/CI/CD pipelines (Jenkins, GitHub Actions).
  • Safe Automation & Guardrails: Experience implementing responsible AI guardrails, deterministic fallback logic, and policy-driven remediation engines.
  • Data & Messaging: Experience with streaming and data platforms (Kafka, Redis, PostgreSQL, Elasticsearch/Vector DBs).

CollabHiring

Why Cisco? 

At Cisco, we’re revolutionizing how data and infrastructure connect and protect organizations in the AI era – and beyond. We’ve been innovating fearlessly for 40 years to create solutions that power how humans and technology work together across the physical and digital worlds. These solutions provide customers with unparalleled security, visibility, and insights across the entire digital footprint.

Fueled by the depth and breadth of our technology, we experiment and create meaningful solutions. Add to that our worldwide network of doers and experts, and you’ll see that the opportunities to grow and build are limitless. We work as a team, collaborating with empathy to make really big things happen on a global scale. Because our solutions are everywhere, our impact is everywhere. 

We are Cisco, and our power starts with you. 

Similar Jobs

3 Minutes Ago
Hybrid
Expert/Leader
Expert/Leader
Fintech • Mobile • Payments • Software • Financial Services
Principal Product Manager responsible for shaping Wise’s Cash & Asset Management and Counterparty Credit Risk product domains. The role owns product strategy, systems, controls, metrics, vendor evaluation, and operating models for managing customer funds and counterparty risk globally. It requires cross-functional leadership across engineering, analytics, and financial stakeholders, with a focus on scalable internal platforms, commercial outcomes, and prudent risk management in a regulated environment.
Entry level
Financial Services
Trade refined oil products by managing client and broker flows, pricing and hedging transactions, developing fundamental and technical trading ideas, and managing complex risk. Responsibilities include marking forward curves, providing market commentary, hedging secondary exposures, sizing positions, supporting client engagement, and maximizing franchise revenue while maintaining strong controls and regulatory standards.
Top Skills: PythonSQL
25 Minutes Ago
Hybrid
Expert/Leader
Expert/Leader
Financial Services
Lead research and development of Transformer-based and time-series foundation models for systematic trading. Responsibilities include pre-training models from scratch on large-scale financial datasets, designing representations and self-supervised objectives, developing distributed training systems, fine-tuning models for trading applications, studying scaling and regime robustness, and evaluating economic performance through simulations and live-trading metrics.
Top Skills: Distributed Training SystemsJaxLarge-Scale Data PipelinesMachine Learning InfrastructureMixed-Precision TrainingModel ServingPyTorchTime-Series Foundation ModelsTransformer Models

What you need to know about the Edinburgh Tech Scene

From traditional pubs and centuries-old universities to sleek shopping malls and glass-paneled office buildings, Edinburgh's architecture reflects its unique blend of history and modernity. But the fusion of past and future isn't just visible in its buildings; it's also shaping the city's economy. Named the United Kingdom's leading technology ecosystem outside of London, Edinburgh plays host to major global companies like Apple and Adobe, as well as a growing number of innovative startups in fields like cybersecurity, finance and healthcare.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account