Senior AI Prompt Engineer
"You'll own the full lifecycle of prompt architecture for real AI products—not just tweaking inputs, but building evaluation frameworks, diagnosing failure modes, and establishing engineering standards across the organization. This is a chance to shape how a company thinks about LLM reliability and quality at scale, working with teams that actually care about moving beyond trial-and-error toward systematic, measurable improvements."
Agency Portfolio is hiring a Senior AI Prompt Engineer on behalf of our client, a company building AI-powered products and services. In this role you will design, test, and refine prompts that get the best possible performance from large language models across real production use cases. You will work closely with product, engineering, and data teams to translate business needs into reliable, well-structured prompting strategies. As a senior member of the team, you will own prompt architecture end to end: developing prompt templates, building evaluation frameworks to measure quality and consistency, and iterating based on results. You will diagnose failure modes such as hallucination, inconsistent formatting, and unsafe outputs, and put guardrails in place to address them. You will also mentor junior team members and help establish prompt engineering standards and documentation across the organization. Responsibilities: Design and optimize prompts for LLM-based features across multiple use cases; build and maintain evaluation and testing pipelines to measure prompt performance objectively; collaborate with engineers to integrate prompts into applications, including retrieval-augmented generation and agent-based workflows; investigate and resolve issues around accuracy, bias, and safety; stay current with new models, techniques, and tooling and recommend improvements; document prompting patterns and share best practices across teams. Requirements: Several years of hands-on experience working with large language models and prompt design in a production or applied setting; strong understanding of how LLMs behave, including context windows, tokenization, and common limitations; experience with techniques such as few-shot prompting, chain-of-thought, and structured output; proficiency in Python for building tooling, tests, and integrations; familiarity with LLM APIs (such as OpenAI, Anthropic, or open-source models) and frameworks for orchestration; a rigorous, data-driven approach to evaluation and iteration; strong written communication skills for documentation and cross-team collaboration. Experience with RAG, vector databases, or agent frameworks is a plus.