You're seeing this page as if you were . The main menu is still yours, though. Exit from immersion
Florian MattanaFM

Florian Mattana

GPU Performance Engineer | CUDA C++ | AI Inference

€850/day
Paris, FR
3-7 years

Average response time: 4 hours

Freelancer profile translated to English.
Back to original language

About Florian

GPU Performance Engineer CUDA C++ | PTX | Inference Optimization

GPU kernel optimization, Nsight profiling, CPU → GPU migration. 5 years of experience (Airbus, DPD Group, Melexis)

GPU Engineer specializing in CUDA C++ kernel optimization and performance profiling. 5 years of experience in aerospace (Airbus), logistics (DPD Group), and semiconductors (Melexis).

My approach: Below PyTorch, not above.
I intervene at the kernel level, where real performance gains are made.

What I can offer you:
→ Development and optimization of CUDA C++ kernels with Nsight Compute profiling, roofline analysis, and production-validated gains (up to 33x speed-up)
→ CPU → GPU migration and production deployment on NVIDIA T4/V100/A10G with numerical validation and cloud deployment
→ GPU performance audit: compute/memory/latency-bound classification, bottleneck identification, actionable optimization plan with quantified metrics
→ GPU mentoring and expertise transfer: team training on profiling, memory optimization patterns, and GPU architecture

Production results:
→ 33x CUDA kernel speed-up (Airbus, V100)
→ GPU utilization 9% → 89% (Airbus)
→ 30% throughput gain, CPU → GPU migration in 6 weeks (DPD Group)
→ 40% reduction in processing time (Datashift, A10G)

In parallel to my assignments, I contribute to the open-source GPU ecosystem (ThunderKittens, model-kernels) and I am developing an FP4 fused attention kernel for Blackwell GPUs in inline PTX.

Technical blog and profiling guide (20,000+ words) at florianmattana.com.

Shall we discuss your project?
Contact me, I respond very quickly.

Skills
CUDA, C++, GPU Computing, NVIDIA, Profiling, Performance Optimization, Nsight Compute, Nsight Systems, Tensor Cores, HPC, CUTLASS, PTX, Inference Optimization, CPU-GPU Migration
  • English

    Native or bilingual

  • French

    Native or bilingual

Can work on-site
Paris (up to 50km), Lyon (up to 50km), Lille (up to 50km), Marseille (up to 50km), Bordeaux (up to 50km)

Experience

  • Open Source | LLM | Inference
    GPU Performance Engineer
    TECH
    January 2026 - Today (7 months)
    Paris, France
    Active contributor to the open-source GPU ecosystem, focused on CUDA inference kernel optimization and performance profiling.

    As a GPU Performance Engineer, I work on:

    → Development of an FP4 fused attention kernel for consumer Blackwell (SM120) in inline PTX — GEMM-softmax-GEMM fused in registers with mma.sync and block scaling UE8M0
    → Bug fixing for compilation and precision issues on existing inference kernels
    → Profiling and performance auditing of real-world GPU kernels with Nsight Compute
    → Writing technical documentation on GPU profiling
    Main contributions:
    → model-kernels: 4 PRs merged — fix for 5 compilation bugs and 2 precision bugs on an INT8 fused attention kernel, max error reduced from 1.69 to 1.37
    → ThunderKittens (Stanford HazyResearch): PR #179 — fix for a narrowing-conversion bug in base-type packing
    → fp4-fused-attention-sm120: FP4 fused attention kernel from scratch for consumer Blackwell GPUs in inline PTX (mma.sync.aligned.mxf8f6f4)
    → CUDA-Kernels: collection of optimized kernels from scratch (GEMM, reduction, prefix scan, softmax, Flash Attention) with full NCU profiling — best GEMM at 58.8% of cuBLAS on RTX 5070 Ti
    → GPU Profiling Guide (20,000+ words) covering Nsight Systems and Nsight Compute end-to-end

    Technical environment: CUDA C++, PTX inline, Tensor Cores, Nsight Compute, Nsight Systems, RTX 5070 Ti (SM120), Git, Linux
    ThunderKittens CUDA Generative AI AI Engineer HPC
  • Melexis
    GPU Performance Engineer
    AUTOMOBILE
    March 2024 - December 2025 (1 year and 9 months)
    Brussels, Belgium
    Melexis is a company specializing in the testing of semiconductor sensors on AWS cloud infrastructure.

    I joined the GPU Compute team to manage the GPU computation pipeline for sensor testing on AWS EC2 g5 (NVIDIA A10G).

    As a GPU Performance Engineer, my responsibilities included:

    → Development and maintenance of the GPU compute pipeline in CUDA C++
    → Multi-precision numerical validation (FP64, FP32, FP16, FP8) with automated CI (cosine similarity ≥ 0.9995)
    → Diagnosis and correction of numerical corruptions (NaN in FP16) via adversarial fuzzing and dynamic range scaling
    → Optimization of host-device transfers with CUDA streams and pinned memory

    I contributed to the following developments:

    → 40% reduction in end-to-end processing time
    → Implementation of an automated multi-precision numerical validation CI gate
    → Improvement of daily throughput through host-device transfer optimization

    Technical environment: CUDA C++, Python, Nsight Compute, Nsight Systems, AWS EC2 (g5, A10G), Docker, GitLab CI, Linux
    CUDA Linux HPC Performance Improvement C++
  • Airbus via Accenture
    GPU Performance Engineer
    AVIATION AND AEROSPACE
    April 2022 - March 2024 (1 year and 11 months)
    Toulouse, France
    Airbus is the world leader in aeronautics and space.

    I joined the satellite inspection team to optimize a CUDA kernel for leak detection on Tesla V100.

    As a GPU Performance Engineer, my responsibilities included:

    → Optimization of the CUDA kernel with Nsight Compute profiling (coalesced memory access, warp divergence elimination)
    → Reduction of shared memory bank conflicts through tile padding and double buffering
    → Production validation of the obtained speed-up
    → Training of 5+ engineers on Nsight Compute profiling workflows

    I contributed to the following developments:

    → 33x kernel speed-up (30 min → < 1 min), validated in production
    GPU utilization from 9% → 89%
    → 40% reduction in shared memory bank conflicts
    → Transition from overnight batch processing to same-day turnaround for satellite inspection jobs

    Technical environment: CUDA C++, Nsight Compute, Nsight Systems, Tesla V100, Python, Linux
    CUDA Linux GPU C/C++ Programming C++ Development

Recommendations

Be the first to recommend Florian

Help this freelancer shine by sharing your experience working together.

These freelancer profiles also match your criteria

AgathaA

Agatha Frydrych

Backend Java Software Engineer

4.7

(3)

2

BaptisteB

Baptiste Duhen

Fullstack developer

4.6

(4)

5

AmedA

Amed Hamou

Senior Lead Developer

4

(2)

7

AudreyA

Audrey Champion

Web developer

4.3

(3)

4

Education

  • Master 2
    Paris 1 - La Sorbonne
    2015
    Finance de marché et gestion des risques

Skill set

Categories