The work
We add AI to the systems businesses already use: extracting fields from documents, classifying incoming information and helping staff find answers in internal knowledge. Your work is to make these features useful, measurable and safe to operate alongside an ERP, a portal or an existing workflow.
What you will do
- Define a narrow use case with the team: the input, expected output, failure cases and the person who checks the result.
- Build document ingestion, extraction and retrieval pipelines using Azure OpenAI and appropriate search or data services.
- Create evaluation datasets and compare changes against a baseline, including incorrect answers, missing information and permission boundaries.
- Integrate AI outputs with APIs and business rules, with validation, traceability and human review where needed.
- Monitor quality, latency and cost, diagnose failures and document how a feature is supported after release.
What you will bring
- Experience building software that uses language models, with an example you can explain beyond the prompt.
- Practical knowledge of retrieval, structured outputs and evaluation, including when a model should abstain or hand work to a person.
- Solid programming skills in C# or Python, familiarity with APIs, Git and testing, and willingness to work with our .NET engineers.
- Care with confidential information, access controls and the difference between a plausible answer and a verified result.
Useful, but not essential
Azure OpenAI, Azure AI Search, document processing, SQL or experience integrating AI into a business application. We value evidence of careful engineering; this is an applied engineering role.
A possible first 90 days
Map a document or knowledge workflow, build a representative evaluation set and deliver a small feature with clear quality criteria and a review path. Then use the findings to improve its behaviour and production monitoring.
How we work
Full time and 100% remote. Client discovery, demonstrations and delivery are online, with no office or client-site attendance required. We agree time-zone overlap during selection and keep experiments, decisions and evaluation results documented.
You will help decide where AI adds value and where a simpler rule is better. Share a project, the way you measured it and a failure that changed your approach. Personal projects are welcome; do not share confidential client data.
