Software Engineer II - ML and Observability
Job Summary:
Technology is at the heart of Disney’s past, present, and future. Disney Entertainment and ESPN Product & Technology is a global organization of engineers, product developers, designers, technologists, data scientists, and more – all working to build and advance the technological backbone for Disney’s media business globally.
The team marries technology with creativity to build world-class products, enhance storytelling, and drive velocity, innovation, and scalability for our businesses. We are Storytellers and Innovators. Creators and Builders. Entertainers and Engineers. We work with every part of The Walt Disney Company’s media portfolio to advance the technological foundation and consumer media touch points serving millions of people around the world.
Here are a few reasons why we think you’d love working here:
- Building the future of Disney’s media: Our Technologists are designing and building the products and platforms that will power our media, advertising, and distribution businesses for years to come.
- Reach, Scale & Impact: More than ever, Disney’s technology and products serve as a signature doorway for fans’ connections with the company’s brands and stories. Disney+. Hulu. ESPN. ABC. ABC News…and many more. These products and brands – and the unmatched stories, storytellers, and events they carry – matter to millions of people globally.
- Innovation: We develop and implement groundbreaking products and techniques that shape industry norms and solve complex and distinctive technical problems.
Product Engineering is a unified team responsible for the engineering of Disney Entertainment & ESPN digital and streaming products and platforms. This includes product engineering, media engineering, quality assurance, engineering behind personalization, commerce, lifecycle, and identity.
The Observability & Insights group ensures that Disney Streaming’s distributed systems are reliable, performant, and transparent. We build ML-powered detection systems, telemetry pipelines, intelligent alerting, and developer experience tooling that enable engineers across the organization to understand system health and take action quickly.
Job Summary:
As a Software Engineer II, you will contribute to building and operating machine learning models and AI-driven systems that enhance the reliability of Disney’s streaming ecosystem. You will work on production ML models - including autoencoders for anomaly detection, statistical threshold systems, and LLM-powered investigation gates - that transform telemetry and signals into automated detection and proactive insights across Disney+, Hulu, and ESPN.
You will participate in the ML lifecycle: feature engineering on time-series data, model training on GPU clusters, real-time inference pipelines, and model improvement. You will partner with engineering and platform teams to embed intelligence into operational workflows, improving system resilience and customer experience at scale.
As a Software Engineer II, you will deliver features end-to-end, participate in design and code reviews, and grow into owning components of our ML systems within a fast-paced, AI-native engineering environment.
Responsibilities and Duties of the Role:
- Contribute to the development of ML models for anomaly detection, including autoencoders, statistical models, and ensemble detection systems that monitor thousands of microservices.
- Develop and improve ML training pipelines using PyTorch on GPU clusters, including feature engineering, model training, calibration and deployment using MLflow.
- Engineer features from time-series telemetry, such as error ratios, latency, and infrastructure metrics, by implementing windowing, normalization, and data quality safeguards.
- Support real-time ML inference systems that run prediction cycles in production, including model serving, detection logic, and alert generation.
- Build AI-driven capabilities using foundation models like Claude and GPT-4 for automated investigation and reasoning over system health signals.
- Create and deploy scalable APIs and services using FastAPI to deliver ML predictions and health insights to engineering teams and operational tooling.
- Partner cross-functionally to integrate ML intelligence into workflows such as incident response and release validation.
- Contribute to model improvement through evaluation, retraining, and threshold tuning.
Basic Qualifications
- 3+ years of professional software engineering experience building, scaling, and maintaining ML-powered backends, data-driven microservices, and production RESTful APIs using FastAPI or Flask.
- Strong hands-on experience in end-to-end ML engineering using PyTorch or TensorFlow, including model architecture selection, feature engineering, training, and evaluation, for example, autoencoders, sequential/time-series models like RNNs/GRUs, anomaly detection, or transformers.
- Proficiency in Python and at least one ML framework, with PyTorch preferred.
- Experience managing model experiments, lineage, and hyperparameter tracking using tools like MLflow or Weights & Biases.
- Practical experience processing, transforming, and querying large-scale telemetry, event, or time-series datasets using PySpark, Pandas, or Databricks.
- Experience with modern development practices including version control (GitHub), containerization (Docker), and cloud-native deployments on AWS (EKS).
- Strong analytical and troubleshooting skills with the ability to iterate rapidly.
- Strong collaboration and communication skills, with the ability to work cross-functionally.
Preferred Qualifications
- Experience with foundation model integration, prompt engineering or evaluation, RAG architectures, or orchestration frameworks like LangChain or LangGraph.
- Familiarity with observability platforms (e.g., Datadog, Grafana, Conviva) and high-volume telemetry data.
- Experience with large-scale data platforms (Databricks, Spark, Snowflake).
Required Education
Bachelor’s degree in Computer Science, Machine Learning, Statistics, Engineering, or equivalent experience.
