Skip to content
Iwana Labs

Engineering

AI Engineer

A senior, hands-on role building agents and AI products that run in production. You will own how agentic systems get evaluated, integrated into existing platforms, and kept reliable once real users depend on them.

Spain
location
Remote
arrangement
Full time
commitment

About

We are an AI consultancy that helps organisations turn data and emerging AI technologies into practical, production-ready solutions.

We are looking for a senior engineer to own the systems side of our AI work: designing agents, integrating them into existing production platforms, and building the evaluation and observability that make them safe to keep changing.

The role suits someone who has already taken agentic or LLM-based products past a demo and then lived with them in production.

First project

Your first engagement will be with a large media platform, working on:

  • Designing and shipping AI agents that automate workflows across content, marketing, analytics, and operations.
  • Integrating agent and LLM capabilities into existing production services, APIs, and data platforms.
  • Building the evaluation, monitoring, and guardrails that let those systems ship changes safely.

The role

  • Design, build, and operate AI agents and tool-calling workflows that run in front of real users.
  • Integrate AI capabilities into existing production systems, including APIs, services, event-driven pipelines, and data platforms.
  • Own the evaluation strategy for AI products: define what good looks like, build the evaluation sets, and gate releases on the results.
  • Build offline and online evaluation, including regression suites, model-graded checks where appropriate, human review loops, and sampling of production output.
  • Design agent architectures with explicit stage boundaries and bounded iteration, choosing simpler approaches when an agentic one does not earn its complexity.
  • Instrument systems for quality, latency, cost, and failure modes, and act on what the telemetry shows.
  • Implement production safeguards such as validation, fallbacks, bounded retries, circuit breakers, access controls, and human approval steps.
  • Version prompts, tools, models, and configuration so that behaviour changes under control.
  • Work with client product, engineering, and data teams to turn business needs into technical designs, and lead delivery through to handover.
  • Set technical standards and build reusable components for AI engineering across engagements.
  • Mentor other engineers and review their work.

You have

  • Typically 5+ years of professional experience in software, machine learning, or AI engineering, including time owning systems in production.
  • Strong proficiency in Python, with experience writing clean, tested, maintainable production code.
  • Demonstrated experience taking LLM or agentic systems past the proof-of-concept stage and operating them afterwards.
  • Deep practical knowledge of AI product evaluation: designing evaluation sets, choosing metrics that track real quality, catching regressions, and gating releases on results.
  • Experience integrating AI capabilities into existing production systems, covering API design, service boundaries, asynchronous or event-driven processing, and data access.
  • Hands-on experience with agent patterns: tool calling, structured outputs, orchestration, retries and fallbacks, and bounded execution.
  • Strong grasp of reliability engineering for non-deterministic systems, including monitoring, alerting, failure modes, latency budgets, and cost control.
  • Solid cloud experience on AWS or Azure, with CI/CD, containers, and infrastructure as code.
  • Ability to evaluate technical approaches on business value, complexity, scalability, reliability, and cost.
  • Strong communication skills and confidence leading technical conversations directly with clients.
  • Comfort operating in ambiguous environments, learning unfamiliar systems, and moving between architecture and implementation.
  • A degree in computer science, engineering, or a related discipline, or equivalent practical experience.

Bonus

  • Experience with evaluation and observability tooling such as LangSmith, Langfuse, or Braintrust.
  • Experience with orchestration frameworks such as LangChain, LangGraph, or Pydantic AI, and with building without them.
  • Experience with multi-provider routing, quota management, and per-request cost attribution.
  • Experience with retrieval systems, including hybrid search, re-ranking, and retrieval-quality evaluation.
  • Experience mentoring engineers or leading small delivery teams.
  • Open-source contributions.
  • Experience in media, publishing, streaming, advertising, e-commerce, or subscription businesses.
  • Experience in a consultancy, agency, startup, or other client-facing delivery environment.

Environment

Foundation models
OpenAI, Anthropic, Alibaba Cloud Qwen, and other commercial or open-weight models.
AI development
Model APIs, prompt and context engineering, structured outputs, tool calling, RAG, embeddings, AI agents, evaluation frameworks, and guardrails.
ML platforms
Amazon SageMaker, AWS Bedrock, Databricks, and MLflow for experimentation, training, model management, serving, and monitoring.
Languages
Python, SQL, pandas, NumPy, scikit-learn, and where appropriate PyTorch or TensorFlow.
Data and processing
Amazon S3, DynamoDB, AWS Glue, Databricks, Spark and PySpark, data lakes, ETL and ELT, and batch or streaming pipelines.
Cloud infrastructure
AWS Lambda, API Gateway, Step Functions, EventBridge, EC2, ECS or EKS, and event-driven or serverless architectures.
Search and retrieval
Elasticsearch, OpenSearch, vector databases, semantic and hybrid search, metadata filtering, and re-ranking.
Engineering and ops
Git, automated testing, CI/CD, Docker, Kubernetes, infrastructure as code with Terraform, CloudFormation or AWS CDK, and monitoring with tools such as CloudWatch.
Security and governance
IAM, secrets management, role-based access, audit logging, data privacy, model governance, and responsible AI controls.

The exact stack varies by client and project, and you are not expected to have used everything on this list. Strong engineering foundations, experience operating production systems, and the ability to learn new platforms quickly matter more than familiarity with any single tool.

First months

  • You understand the client's platform, data, and the constraints the AI has to live inside.
  • You have shipped an agent or AI workflow into production with evaluation and monitoring already in place.
  • You have established the evaluation approach the team uses to decide whether a change is safe to ship.
  • You have built trusted relationships with client engineering, product, and data stakeholders.
  • You are setting the technical standards other engineers on the engagement follow.

Perks

  • Remote-first anywhere in Spain. The team gets together in person a few times a year, and you work wherever you work best the rest of the time.
  • 22 days of paid time off, plus public holidays.
  • Ticket Restaurant meal allowance and private health insurance.
  • A learning budget for conferences, courses, and books.
  • Access to the latest AI agents and tooling to help you do the work.
  • A Mac, and the rest of the setup you need to work comfortably.

Why join

  • Own AI systems end to end, from architecture through production operation.
  • Work on agents and AI products used by large and engaged audiences.
  • Work directly with clients and influence both technical architecture and product direction.
  • Help shape the reusable tools, standards, and delivery practices of a growing AI consultancy.

Process

  1. Intro call

    Thirty minutes with a founder. What you have built, what you are looking for, and what the work here actually involves.

  2. Take-home assessment

    A practical exercise close to the work we do, completed in your own time. We design it to take about three hours, and we will not ask for more than that.

  3. Technical deep dive

    We go through your solution together: the decisions you made, the trade-offs you weighed, and what you would change with more time.

  4. Feedback or offer

    We come back to you either way, with feedback you can use.

The whole process usually takes two to three weeks from the intro call, and we work around your schedule.

How to apply

Email us your CV and something you have built: a repository, a system you shipped, or a write-up. We read every application ourselves and reply either way.

Apply for AI Engineer