Loading...

ML Systems Engineer — Inference Acceleration

  • Company: Get A Job.ai
  • Location: Paris Offices
  • Salary: Pay not listed

Website Get A Job.ai

Represented by Get A Job.ai

About This Opportunity

We are representing a confidential deep-tech semiconductor organization at the forefront of next-generation processor architecture. Our client is pioneering novel approaches to AI acceleration through proprietary hardware technology, backed by leading investors and industry veterans. We're seeking an experienced ML Systems Engineer to join their Paris offices and optimize inference workloads on their custom accelerator platform.

Responsibilities

  • Analyze modern AI workloads and identify kernel-level, runtime, memory, and system-level bottlenecks on custom accelerator hardware
  • Develop and optimize custom kernels, fused operators, and execution strategies to maximize device utilization
  • Design efficient mappings of models and operators across multiple accelerator devices, including communication and synchronization strategies
  • Implement inference-serving techniques such as continuous batching, paged KV caches, prefix/context caching, chunked prefill, and prefill/decode interleaving or disaggregation
  • Build profiling, benchmarking, and performance-analysis infrastructure spanning kernels, full models, and serving workloads
  • Collaborate closely with hardware, compiler, and runtime teams to co-design software abstractions and influence future hardware features based on real model workloads

What We're Looking For

  • Strong experience in high-performance ML inference, GPU/accelerator programming, or ML systems engineering
  • Deep understanding of computer architecture, accelerator/GPU execution models, memory hierarchies, parallelism, and performance bottlenecks
  • Experience developing and optimizing custom kernels using CUDA, Triton, ROCm/HIP, or equivalent low-level programming environments
  • Experience with operator fusion, tiling, scheduling, data movement optimization, graph execution, and profiling of compute- and memory-bound workloads
  • Strong understanding of distributed model execution, including tensor, pipeline, sequence, and/or expert parallelism and communication/computation overlap
  • Hands-on experience with modern inference-serving systems such as vLLM, SGLang, TensorRT-LLM, or equivalent, including KV-cache management, continuous batching, paged attention, and prefill/decode scheduling
  • Strong C++ and Python skills, and comfort working on a custom accelerator stack where compiler, runtime, kernels, and abstractions are actively being developed
  • Exposure to or experience with MLIR and MLIR dialects is a strong plus
  • Proficient English language skills

How We Work With You

Candidates apply directly through Get A Job.ai. Our recruiting team will conduct an initial screening to understand your background and ensure alignment with the role requirements. Qualified candidates are then presented to our client for their review and interview process. Please note that you should not contact the client organization directly—all communication and coordination will be handled through our talent team at Get A Job.ai.

Compensation and Benefits

Our client offers competitive cash compensation based on location, experience, and internal equity. The majority of full-time offers include a meaningful stock option plan. Benefits include healthcare coverage with family-friendly options, pension contributions, professional development support, and 25 days of PTO in addition to public holidays. This role offers ownership of a key technical domain with significant growth opportunities based on performance and individual drive.

Equal Employment Opportunity: Get A Job.ai is committed to providing equal employment opportunities to all qualified applicants without regard to race, color, religion, sex, national origin, age, disability, or any other protected characteristic. We value diversity and encourage applications from all qualified candidates.

Apply with Get A Job.ai

A recruiter will review your profile and submit you to the client. Do not contact the client directly.

More options

Apply with Get A Job.ai

Apply through Get A Job.ai. A recruiter will review your profile and submit you. Do not contact the client directly.

Apply through Get A Job.ai. A recruiter will review your profile and submit you.

Terms used in this posting

PTO
Paid Time Off — vacation, personal, or sick days you can take while still being paid.
equity
Ownership stake in the company, usually in the form of stock options or RSUs, offered in addition to salary.

Working in Paris offices

    Job details above are provided by the employer/source. The sections on this page are compiled from public data sources with AI assistance.

    Accommodations: if you need a workplace accommodation to apply for or perform this job, see ADA.gov or EEOC.gov for guidance on your rights and how to request one.

    Add application deadline to calendar

    Listing facts

    • Role ML Systems Engineer — Inference Acceleration
    • Employer Get A Job.ai
    • Location Paris Offices
    • Type Full Time
    • Pay (from listing) Pay not listed
    • Posted September 13, 2026
    • Apply by October 13, 2026
    • Overview Full job description on this page (494 words)

    Facts above come from this job record on Get A Job.AI — not copied from third-party review sites.

    Limited public data for this employer

    We only show facts we can ground in public sources (Wikidata, O*NET, news/discussion links, or this listing). We do not invent Glassdoor-style ratings, salaries, or testimonials when data is thin. Use the listing facts, occupation context, and related openings below while we continue researching.

    Explore related openings

    Keep exploring on Get A Job.ai

    Not quite the right fit? Your next opportunity is a click away.

    Hiring instead? Post a job and reach candidates searching right now.

    ML Systems Engineer — Inference Acceleration Get A Job.ai · Paris Offices