Senior Machine Learning Infrastructure Engineer
Build the systems that take AI from experiment to production.
We usually respond within a day
About eComID
eComID is building the Shopping Passport for modern commerce - a shopper context layer that enables brands to understand and personalize every shopper from their very first visit. By combining AI-powered sizing, conversational search and intelligent product discovery, we’re making online shopping smarter, more personal and less wasteful.
Launched in Stockholm in 2024, eComID is already operating at scale, reaching more than 20 million shoppers every month. We recently raised a $17.4 million seed round - the largest in European fashion-tech history and the second-largest globally.
We’re backed by H&M Group, Stadium, leading international VCs and some of Europe’s most successful industry leaders and technology founders, including Helena Helmersson, Sebastian Knutsson, Maria Raga and Alan Mamedi.
Now, we’re bringing together exceptional people with warm hearts to build a world-class fashion-tech company from Stockholm. If you want to solve meaningful problems, move fast and help shape how the world shops, this is the place to do it.
What you will work on
We are looking for a Senior MLOps Engineer to own and evolve the systems that take our models from development into production.
You will sit at the intersection of machine learning, software engineering, and infrastructure. Your job is to make it easy for our AI engineers to build, deploy, evaluate, monitor, and improve models while ensuring those systems remain reliable, scalable, observable, and efficient in production.
This is not a role focused only on maintaining infrastructure. You will work directly with the engineers building our models and products, helping shape both the platform and the way we develop AI systems at eComID.
You will help build and own the infrastructure and tooling behind our production AI systems, including:
Building and improving model training, evaluation, and deployment pipelines
Designing reliable production systems for real-time and batch model inference
Improving how models move from experimentation to production
Building tooling that makes it easier and safer for AI engineers to ship new models
Model versioning, reproducibility, rollbacks, and release strategies
Monitoring model performance, latency, errors, resource usage, and production health
Designing automated evaluation and validation before and after deployment
Scaling inference workloads while balancing latency, reliability, and infrastructure cost
Managing feature and model dependencies across training and inference
Improving CI/CD workflows for machine learning services and pipelines
Building observability across models, pipelines, services, and infrastructure
Designing systems for experimentation, shadow deployments, A/B testing, and gradual rollouts
Making our AI platform simpler, faster, and more reliable as both our team and product surface grow
You will work closely with AI engineers to understand where the development process is slow, fragile, or repetitive and turn those problems into better infrastructure and tooling.
What we are looking for
We are looking for someone with strong software and infrastructure fundamentals who has experience operating machine learning systems in production.
You should have experience with several of the following:
Building and operating production ML infrastructure
Deploying and serving machine learning models at scale
Designing CI/CD systems for ML or backend services
Cloud infrastructure, ideally GCP
Containerised workloads and orchestration
Python in production environments
Infrastructure as code and automated environment management
Monitoring, observability, alerting, and production debugging
Model versioning, experiment tracking, and reproducibility
Training and inference pipelines
Distributed workloads and asynchronous processing
Real-time, low-latency inference systems
GPU workloads and inference optimisation
Reliability engineering and production incident handling
Designing internal platforms or developer tooling
Experience working directly with ML engineers or data scientists is important. You should understand enough about the ML lifecycle to identify where infrastructure can improve experimentation, deployment, evaluation, and production reliability.
We care more about strong engineering judgement and experience operating real systems than whether you have used our exact stack before.
Our current tech stack
Today our core stack includes:
Languages: Python, Go, TypeScript
Services/libraries: GCP, BigQuery, ClickHouse, PostgreSQL, Redis, Apache Beam, Kubeflow, Feast, PyTorch, Typesense, Milvus, ConnectRPC
We are not attached to technologies for their own sake and are open to introducing better tools where they make sense.
What we value
We are looking for someone who:
Takes ownership of the systems they build
Cares about reliability and understands what it means to operate systems in production
Enjoys working directly with AI engineers and product teams
Thinks carefully about failure modes, observability, reproducibility, and operational complexity
Can reason about trade-offs between latency, throughput, reliability, cost, and developer experience
Will question existing architecture and propose better approaches
Prefers simple, robust systems over unnecessary complexity
Automates repetitive work rather than accepting it as part of the process
Treats infrastructure as a product used by other engineers
Cares about making the path from experiment to production fast and predictable
Is comfortable debugging problems across application code, infrastructure, networking, data, and models
Thinks about the full lifecycle of a model, not only how to deploy it
Can build abstractions without hiding the underlying system when things go wrong
Knows when to introduce platform capabilities and when a simpler solution is enough
What we offer
A best-in-class team with warm hearts and no ego. We care deeply about the quality of our work, while staying kind, humble, and easy to work with. The goal is always to build the best product possible, not to win internal arguments.
A direct connection to product. You will work with product every day, understand why we are building something, help shape the solution, and own the delivery from idea to production.
Real ownership. You will have meaningful influence over our ML platform, infrastructure, tooling, and engineering decisions. We expect senior engineers to improve the way we work, not just maintain what already exists.
Work directly with the AI team. You will sit close to the engineers building and experimenting with our models, helping turn research and prototypes into reliable production systems.
Build infrastructure that directly changes how fast we can innovate. Improvements you make to deployment, evaluation, observability, and tooling will directly affect how quickly the team can experiment and ship.
Interesting technical problems with real users behind them. You will work on production AI systems where latency, reliability, scalability, and model quality all matter.
Freedom to choose the right tools. We have a strong existing stack, but no attachment to technology for its own sake. If there is a better way to solve a problem, we want to hear it.
High standards without unnecessary process. We care about reliability, thoughtful engineering, and maintainable systems while keeping teams small, communication direct, and decision-making fast.
The chance to shape the next stage of the company. We are still early enough that the platform, principles, and engineering practices you establish now will have a lasting impact on how eComID builds and operates AI.
Benefits
Holiday: 25 days of paid holiday per year
Health allowance: 5,000 SEK wellness allowance per year
Office: Work from our office in central Stockholm
Office perks: Weekly breakfasts and snacks throughout the day
Tech: Your choice of computer, phone, and phone contract
Learning & growth: Annual budget for courses, books, and conferences
Insurance: Extensive coverage, including private healthcare insurance
Pension: Private pension insurance via SPP
- Department
- AI
- Locations
- eComID HQ
- Remote status
- Hybrid
About eComID
eComID is building the Shopping Passport for modern commerce - a shopper context layer that enables brands to understand and personalize every shopper from their very first visit. By combining AI-powered sizing, conversational search and intelligent product discovery, we’re making online shopping smarter, more personal and less wasteful.
Launched in Stockholm in 2024, eComID is already operating at scale, reaching more than 20 million shoppers every month. We recently raised a $17.4 million seed round - the largest in European fashion-tech history and the second-largest globally.
We’re backed by H&M Group, Stadium, leading international VCs and some of Europe’s most successful industry leaders and technology founders, including Helena Helmersson, Sebastian Knutsson, Maria Raga and Alan Mamedi.
Now, we’re bringing together exceptional people with warm hearts to build a world-class fashion-tech company from Stockholm. If you want to solve meaningful problems, move fast and help shape how the world shops, this is the place to do it.