- Company: Get A Job.ai
- Location: France
- Salary: Pay not listed
Website Get A Job.ai
Represented by Get A Job.ai
About This Opportunity
We are representing a confidential client in the family technology space, based in France. They are seeking an AI Engineer to own and optimize their self-hosted inference infrastructure while also shipping user-facing AI features across multiple consumer applications.
This is a hands-on technical role where you'll run GPU inference workloads on dedicated hardware, build internal AI platforms, and deliver product-facing agents that help families coordinate their daily lives. You'll work across the full stack—from low-level inference optimization to app-layer feature development.
Responsibilities
- Run and optimize self-hosted inference infrastructure: Deploy and tune serving stacks (vLLM, SGLang, TensorRT-LLM) on dedicated GPU hardware for high throughput and low latency
- Drive aggressive performance optimization: Implement tensor parallelism, quantization (FP8, AWQ, GPTQ), KV-cache and prefix caching, continuous batching, speculative decoding, and concurrency tuning
- Manage multi-model serving: Balance internal workloads and latency-sensitive product traffic through multi-LoRA, routing, and request scheduling
- Improve efficiency and observability: Instrument performance metrics, improve GPU utilization, and surface technical tradeoffs to inform decision-making
- Ship AI features and proactive agents: Build in-app agent layers including proactive nudges, smart suggestions, summarization, drafting, scheduling, and task automation
- Build supporting infrastructure: Create tools, memory systems, orchestration, guardrails, and evaluation harnesses integrated with production APIs
- Rapid prototyping: Work collaboratively to quickly test ideas, including building rough UIs when needed, then harden successful features
What We're Looking For
Required qualifications:
- 5+ years shipping production software with meaningful applied AI or ML experience
- Demonstrated experience running and optimizing self-hosted LLMs on dedicated multi-GPU hardware
- Hands-on expertise with serving stacks such as vLLM, SGLang, or TensorRT-LLM
- Deep knowledge of inference optimization techniques: tensor parallelism, quantization, batching strategies, and KV cache management
- Proven track record optimizing latency, throughput, and GPU utilization
- Strong Python engineering skills with full-stack capabilities
- Experience with agent frameworks (Claude Agent SDK, LangGraph, or similar), LLM APIs, embeddings, and RAG
- Proficiency with AWS, Docker, CI/CD, monitoring, and observability tools
- Experience building internal platforms and tooling that other teams depend on
Ideal candidate attributes:
- Full-stack mindset—you want to work on app-layer features, not just infrastructure
- Performance-focused approach treating latency and efficiency as core engineering concerns
- Comfortable with rapid prototyping and modern AI-first development tools
- Motivated by building technology that makes a meaningful difference in people's lives
- Bonus: Experience with Slack apps, MCP, or agent orchestration at team scale
How We Work With You
When you apply through Get A Job.ai, our talent team will review your profile and conduct an initial screening. If there's a strong match, we'll submit your candidacy directly to our client and guide you through their interview process.
Please apply exclusively through Get A Job.ai—do not contact the client directly, as this is a confidential search.
Pay
Compensation details will be discussed during the screening process based on experience and qualifications.
Equal Opportunity: Get A Job.ai is committed to inclusive recruiting practices. We welcome candidates from all backgrounds and work to ensure fair consideration throughout the hiring process.
Apply with Get A Job.ai
A recruiter will review your profile and submit you to the client. Do not contact the client directly.
Apply through Get A Job.ai. A recruiter will review your profile and submit you. Do not contact the client directly.
Apply through Get A Job.ai. A recruiter will review your profile and submit you.
Local insights for this role are preparing — this section updates automatically in a few seconds (or refresh).
Listing facts
- Role AI Engineer
- Employer Get A Job.ai
- Location France
- Type Full Time
- Pay (from listing) Pay not listed
- Posted September 1, 2026
- Apply by October 2, 2026
- Country France
- Overview Full job description on this page (492 words)
Facts above come from this job record on Get A Job.AI — not copied from third-party review sites.
Typical work in AI engineer
Independent occupational context from O*NET (U.S. public-domain labor data). This is about the occupation, not a rewrite of this employer's posting.
- Analyze problems to develop solutions involving computer hardware and software.
- Apply theoretical expertise and innovation to create or apply new technology, such as adapting principles for applying computers to new uses.
- Assign or schedule tasks to meet work priorities and goals.
- Meet with managers, vendors, and others to solicit cooperation and resolve problems.
- Design computers and the software that runs them.
- Conduct logical analyses of business, scientific, engineering, and other technical problems, formulating mathematical models of problems for solution by computers.
Limited public data for this employer
We only show facts we can ground in public sources (Wikidata, O*NET, news/discussion links, or this listing). We do not invent Glassdoor-style ratings, salaries, or testimonials when data is thin. Use the listing facts, occupation context, and related openings below while we continue researching.
Explore related openings
Occupation family: AI engineer
Keep exploring on Get A Job.ai
Not quite the right fit? Your next opportunity is a click away.
- Browse all jobs
- More jobs by category
- Remote jobs you can do from anywhere
- Research typical pay for this role
- Set a job alert so new matches reach you first
- Upload your resume to apply faster
Hiring instead? Post a job and reach candidates searching right now.
