Ankon Mukherjee

I build ML systems and the backends around them.

Final-year B.Tech student in AI & Data Science in Mumbai. I care about the part that usually gets skipped: results you can reproduce, and failures that make noise instead of hiding.

Looking for a semester-long internship in 2027 and full-time roles in ML systems, backend engineering or GPU compute.

NeRF render of a Lego bulldozer trained on clean views: fully intact and sharp.

How much poisoned training data does it take to erase an object from a NeRF?

Clean training views: target‑region PSNR 23.6 dB.

Share of target views poisoned

Figure 1. The same NeRF, retrained with 0, 20 and 50% of its target views hard‑erased: the bulldozer painted out with flat mid-gray, RGB (128, 128, 128), which this panel is based on. PSNR is measured inside the bulldozer’s mask and averaged over four evaluation views; the image is one of them. Three runs, not a curve.

Work

Targeted data poisoning of neural radiance fields

Independent research Sep 2026 onward Python, PyTorch, Blender, CUDA

Problem. A NeRF learns a 3D scene from a set of photos. If an attacker can edit some of those photos, can they remove one object from the scene, and what does that do to everything else? To measure it I built a Blender pipeline for a procedural scene and camera rig, a scripted compositor that erases the target from selected views, retraining, and evaluation of the target region and the rest of the frame separately, on a held-out view set that was frozen and checksummed before the first poisoned run.

Key decision. The methodology was locked before any results existed, and every non-trivial choice since is logged with the alternatives I rejected. That discipline paid off before the main sweep: it caught four silent config defects, including every poisoned condition quietly training on clean data, before any compute was spent on them.

Result. The proof of concept passed: target‑region PSNR fell from 23.6 to 19.4 to 11.6 dB as the poisoned share went from 0 to 20 to 50% (Figure 1). Next is an 8-condition, 3-seed sweep at 200,000 iterations per run.

ActAudit

EU AI Act risk classifier Python, Gemini API, Streamlit

Problem. Give it a GitHub repository or a product description, and it places the AI system in an EU AI Act risk tier, with UNESCO and IEEE ethics flags. Asking an LLM whether a system is high-risk gives an answer nobody can audit, and it can change between runs.

Key decision. The LLM only extracts facts into a fixed schema, quoting the text as evidence for each one. An ordered table of plain Python rules makes the decision: the first match wins, and the same facts always give the same tier, with the rule and legal provision that decided it.

Result. Deployed and public, with 102 tests.

LayerWise

Explainable transfer-learning advisor In progress Python, PyTorch, FastAPI

Problem. Fine-tuning a pretrained model means a lot of judgment calls: which model, which layers to freeze, what learning rate. LayerWise is meant to look at an image dataset and make those calls with a reason attached to each one.

Key decision. The engine runs independently of the API, so every decision can be tested without a server.

Result so far. A dataset analyzer that computes per-class statistics over sampled images, and a domain detector that tells natural photos from medical, satellite and other imagery, with a confidence score and a readable chain of evidence. The recommender itself is not built yet; it is next.

Right now

Besides the poisoning sweep, I’m building a roofline profiling study of a small transformer: deriving FLOP and byte counts by hand, then checking them against what the GPU actually does, in fp32 and fp16, to find where each layer flips from compute-bound to memory-bound.

Experience and education

2023 – 2027 (expected)

B.Tech in AI & Data Science, Honours in Applied Cybersecurity

KJ Somaiya College of Engineering, Mumbai. CGPA 9.22 of 10.

Sep 2023 – Sep 2026

Orion Racing India, Formula Student

Aerodynamics lead, Jan 2025 – Sep 2026 Aerodynamics engineer, Sep 2023 – Jan 2025

Led a team of 10+ engineers across design, simulation and manufacturing, and improved the correlation between our CFD simulations (SimScale) and on-track performance. Drove the car to first place in Skidpad among 25+ teams.

Mar – May 2025

Red team analyst intern, DeepCytes

UK, remote

Analyzed macOS and iOS apps using static analysis and reverse engineering in Ghidra. Contributed to data ingestion pipelines and backend APIs for threat-intelligence workflows, and helped automate security analysis workflows.

Contact

Email is the fastest way to reach me. I’m also on LinkedIn and GitHub, and my résumé is a one-page PDF.