Rebuild the held-back evaluation set and tell us honestly where we are weak.
OPEN ROLE / Matching and screening / 02 OF 05
Applied ML Engineer
Make the scoring good enough that a hiring manager argues with the criteria instead of the software.
THE ROLE / WHAT YOU OWN
About this role.
Every rank in Atsifly cites the exact lines that produced it. That constraint is the product: it is what lets a debrief check the work and a rejected candidate be given a real reason.
You would own matching, screening and the evaluation harness that decides whether a change is actually better rather than merely newer.
THE RAMP / NO WARM-UP LAPS
Your first 90 days.
Every role here starts with real work on the live product. This is the ramp we will agree on together, and the pace we hire for.
Ship a scoring change that measurably improves recall on adjacent experience.
Own the release gate: nothing reaches a live workspace without passing it.
RESPONSIBILITIES / THE WORK
What you will do.
- 01
Own candidate matching, screening extraction and score explanation
- 02
Build and maintain the held-back evaluation set and its harness
- 03
Design the bias and fairness checks that gate every release
- 04
Keep latency and cost inside the budget a live desk can afford
- 05
Write down what the model cannot do, and make the product say so
TECH / THE STACK
What you will work with.
Python and FastAPI on the backend, Vue on the front, Postgres and Redis underneath, running on Azure.
MODELLING
- Python
- structured output
- retrieval
- evaluation harnesses
DATA
- Postgres
- pgvector
- Celery
- Redis
PLATFORM
- Azure
- Docker
- observability
QUALIFICATIONS / THE BAR
What you bring.
Must have
- Shipped applied machine learning into a product people pay for
- Strong opinions about evaluation, held-back sets and regression
- The discipline to report a negative result clearly
- Comfort working with language models as components rather than magic
Bonus points
- Fairness or bias testing under a regulatory constraint
- Information extraction from messy documents
- Experience in hiring, credit or another decision-heavy domain
THE CULTURE / DAY TO DAY
How we work.
EVIDENCE OVER OPINION
A scored shortlist against real past hires settles the debate. The best argument in the room is a measurement.
WEEKS, NOT QUARTERS
Work ships to real desks fast. You will watch a recruiter use what you built this month.
SMALL AND SENIOR
No layers, no committees. A handful of hires, direct access to the founders.
REAL DESKS
We sit with recruiters on live roles. If it survives a Friday afternoon, it works.
WHAT WE OFFER / STRAIGHT TERMS
Own part of what you build.
A senior package with founding-team equity and profit sharing on top. We are a startup: the upside is real and the ownership is yours.
- A senior package with the same straight terms for every open role
- Founding-team equity: you own a piece of what you build
- Profit sharing once the product is earning
- Senior ownership of a whole layer of the product, with direct access to both founders
- A hybrid base, plus time sitting with recruiters on live roles
- The hardware and tools you need, without a procurement fight
APPLY / Matching and screening
Tell us what you would build first.
A scoring system that survives a debrief, an evaluation harness the team trusts, and a bias gate nobody argues about. Send a short note with An evaluation you designed, and what it told you that you did not want to hear.. A conversation with a founder follows within days.
Apply as Applied ML Engineer