magnifier icon

Principal Software Engineer, AI Platform Engineering

Saviynt

Saviynt

Internet Software

Milpitas, CA - USA

Entry Level

Experteer Overview

In this role you shape the AI platform's data architecture, governing training data flow and ensuring tenant isolation and privacy. You will set standards used by ML engineers and scientists, influencing end-to-end data lifecycle from source to model. You’ll own infrastructure for data lakes, pipelines, schema management, and multi-tenant safety, enabling scalable, compliant AI workloads. This position offers impact across a large-scale SaaS platform and collaboration with cross-functional teams to advance secure, reliable AI capabilities.

Compensation / Benefits

  • competitive compensation
  • growth opportunities
  • reliable, enterprise-grade platform
  • collaborative, reliability-focused culture

Responsibilities

  • Define and govern data lake architecture on GCS (raw, silver, gold; CMEK; lifecycle)
  • Build and maintain batch pipelines (Spark on Dataproc; Iceberg maintenance; daily S3GCS sync)
  • Develop streaming pipelines (Apache Beam on Dataflow; exactly-once semantics; PII gating)
  • Manage schema registry (Avro/Protobuf versioning; migration playbooks)
  • Oversee orchestration (Flyte; domain isolation; retry policies; DataCatalog memoization)
  • Enforce multi-tenant data architecture (per-tenant isolation; quotas; contamination checks)
  • Develop Data Anonymizer and Data Labeler microservices for PII stripping and labeling
  • Operate feature store (offline/online; consistency; PIT joins) and vector store (Pgvector; Qdrant)
  • Design RAG data pipelines (embedding generation; upsert; freshness SLAs)
  • Expose data platform services via API (gRPC/HTTPS; mTLS)
  • Create synthetic data pipelines for dev/staging
  • Implement data quality gates (Great Expectations/dbt) as deployment-time checks

Key requirements

  • 8+ years of data engineering at production scale
  • Proven impact leading platform-wide standards or major migrations
  • Data lake ownership and end-to-end operation experience
  • Deep Spark (PySpark/Scala) expertise; Iceberg/Delta Lake maintenance
  • Hands-on Beam/Dataflow experience (windowing; exactly-once)
  • Schema registry experience (Protobuf/Avro compatibility; migrations)
  • Production orchestration at scale (Flyte; Kubeflow Pipelines; Airflow; Prefect)
  • Multi-tenant architecture with strict isolation
  • Feature store operations (Feast or Tecton); point-in-time joins; online/offline consistency
  • Vector databases in production (Pgvector, Qdrant); embedding upserts; index strategies
  • RAG fundamentals: chunking, embeddings, retrieval quality; context freshness
  • API transport (gRPC, HTTPS/mTLS); proto contracts
  • Bachelor in Computer Science
  • equivalent practical experience

Description

In this role you shape the AI platform's data architecture, governing training data flow and ensuring tenant isolation and privacy. You will…
For members onlyMobile Experteer Ad

Take your next career step

  • 1M+ top positions worldwide with salary benchmarks

  • Be discreetly found and contacted by headhunters

  • Exclusively for senior-level professionals and executives

Already a member?