Forte Group is hiring AI Product Engineer to build AI-native applications for enterprise clients across the US and Europe. We do not just build prototypes; we ship production-grade LLM features, retrieval systems, and complex agentic workflows. Our recent projects include multi-tenant clinical platforms that slashed validation times and revenue optimization systems processing massive datasets. These systems handle real data, strict compliance, and real consequences when they break.
What You'll Do
Architect RAG & LLM Pipelines: Handle document ingestion, advanced chunking strategies, vector store integration, and re-ranking for complex, multi-format data sources.
Develop Agentic Workflows: Build multi-agent coordination using tool calling, Human-in-the-Loop (HITL) checkpoints, and MCP endpoints for strict data governance.
Build Evaluation Infrastructure: Implement automated test sets, field-level accuracy metrics, and CI/CD regression detection to catch model degradation early.
Own Production Operations: Manage latency optimization, token budgeting, cost tracking, and incident response for non-deterministic AI systems.
Core Requirements
Experience & Fundamentals: 4+ years in
software engineering with 2+ years specifically building and scaling AI-powered features in production.
Backend & Architecture: Strong proficiency in a modern backend language (Python, C#/.NET, Java/Kotlin, TypeScript, or Go) and related web frameworks. Deep understanding of async/concurrent patterns, streaming (SSE/WebSockets), and batch processing.
LLMs & Prompt Engineering: Expertise with OpenAI, Anthropic, and open-source models (Ollama, vLLM). Experience with systems-level prompting, structured output, chain-of-thought, function calling, and multi-model routing (fallback chains, tiering by cost/complexity).
RAG & Retrieval: Hands-on with Vector databases (Pinecone, pgvector, Weaviate, etc.), hybrid search (vector + BM25), and re-ranking. You understand embedding models and the nuanced tradeoffs between fixed, semantic, and sentence-based chunking strategies.
Document Intelligence: Proven ability to build pipelines for unstructured data ingestion, multi-format handling (PDFs, Excel, images), and OCR integration (Azure Document Intelligence or similar).
Agents & Orchestration: Experience with frameworks like LangChain, LangGraph, CrewAI, or building custom orchestration engines. Familiarity with MCP (Model Context Protocol) and implementing safe tool-use patterns (sandboxed execution, approval gates).
AI Ops & Observability: Strong background in AI CI/CD, evaluation frameworks (LangSmith, Braintrust, Ragas), and tracing multi-step AI pipelines. You know how to track model drift, latency, and per-query costs, and how to implement strict data governance (PII handling) in regulated environments.
What Sets You Apart
You can walk us through an AI feature you shipped to production and tell us what went wrong. When you talk about RAG, you talk about chunking tradeoffs and retrieval accuracy, not just that you used a vector database. You have dealt with latency at scale, cost surprises, inputs the model was not trained for, and explaining AI behavior to stakeholders who care about results, not architecture.
You build evaluation into your systems from day one because you know firsthand the consequences of launching without guardrails.