Skip to content
View Thehunk1206's full-sized avatar
🎯
Focusing
🎯
Focusing

Block or report Thehunk1206

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Thehunk1206/README.md

Generative AI, computer vision, and robot learning banner

Tauhid Khan

AI Research Engineer | Generative AI | Computer Vision | Robot Learning

I build research-driven machine learning systems, from model design and large-scale training to efficient production deployment.

LinkedIn Google Scholar IsoFM on OpenReview Email


πŸ‘‹ About Me

I build research-driven machine learning systems that move from model design and distributed training to optimized production inference.

I am an AI Research Engineer with 4+ years of experience across generative AI, computer vision, multimodal learning, and efficient inference. At Fynd, I develop image and video systems involving Flow Matching, diffusion models, super-resolution, model distillation, distributed PyTorch, and TensorRT.

Previously, at Wobot.ai, I built large-scale video analytics systems for detection, multi-object tracking, and cross-camera association. My current research connects generative vision with embodied intelligence: learning representations, trajectories, and policies that are accurate, efficient, and deployable.

4+ years of experience 20K daily active users 1,000+ cameras 95.5% robot-learning success rate

πŸ”¬ Research And Engineering Focus

Generative AI and Flow Matching research visualization

✨ Generative Vision

Flow Matching Β· Diffusion Β· Image and video generation Β· Super-resolution Β· One-step distillation
Computer vision production pipeline visualization

πŸ‘οΈ Computer Vision

Detection Β· Segmentation Β· Multi-object tracking Β· Restoration Β· Production GPU inference
Robot learning and manipulation policy visualization

πŸ€– Robot Learning

VLA Β· Visuomotor policies Β· Imitation learning Β· Reinforcement learning Β· Sim2Real

πŸ“Š Selected Impact

πŸš€ 4x

Flow Matching Super-Resolution
Distilled into a single-step generator

⚑ 40-45%

Lower Inference Latency
TensorRT-optimized deployment

🌐 20K

Daily Active Users
Production image and video services

🎬 56%

Lower Video Latency
Up to 52% lower compute cost

πŸ“Ή 1,000+

Camera Deployments
Across more than 250 locations

πŸ€– 95.5%

StackCube-v1 Success
Flow Matching Transformer + ViT policy

High-resolution inference: 2500x2500 inputs processed in approximately 4-5 seconds on one NVIDIA L4 GPU.

πŸ“„ Publication

🧭 Isokinetic Flow Matching for Pathwise Straightening of Generative Flows

Accepted at the ICML 2026 Workshop on Structured Probabilistic Inference and Generative Modeling (SPIGM).

A geometry-aware Flow Matching formulation for straighter generative transport paths and more efficient generation.

Read IsoFM on OpenReview

πŸ’Ό Experience

🧠 ML Research Engineer · Fynd

Sep 2023 - Present Β· Flow Matching super-resolution, single-step distillation, SDXL and FLUX training, video segmentation and inpainting, controllable generation, distributed PyTorch, and TensorRT deployment.

πŸ‘οΈ Computer Vision Engineer Β· Wobot.ai

Feb 2022 - Sep 2023 Β· Production detection and tracking, multi-camera analytics, CPU and GPU pipeline optimization, Docker, and NVIDIA Triton deployment.

πŸš€ Featured Projects

πŸ€– mini-pi0

Multimodal robot-learning framework using ManiSkill, MuJoCo, wrist-camera observations, and visuomotor Flow Matching policies.

95.5% StackCube-v1 success Β· VLA Β· IL Β· RL

View benchmarks β†’

Conditional Flow Matching experiments for image and video generation with optimal-transport paths and ODE sampling.

13+ stars Β· PyTorch Β· Generative Modeling

🎞️ VideoGen MeanFlow

Complete video-generation training and inference pipeline using MeanFlow with a Diffusion Transformer.

Video Generation Β· DiT Β· Flow Models

πŸŒ™ Zero-DCE

TensorFlow implementation of zero-reference low-light image enhancement.

49+ stars Β· 8+ forks

More applied ML work: 🩺 PraNet Polyp Segmentation Β· πŸŽ™οΈ TinyML Audio Classification

🦾 Current Build

I am building an SO-ARM101 from individual components rather than using a pre-assembled system. The goal is to connect simulation-based policy development with physical data collection, evaluation, and learned-policy deployment. Hardware integration and real-world policy testing are currently in progress.

🧰 Technical Skills

🧠 Core Machine Learning

Python PyTorch TensorFlow OpenCV Linux

Representation Learning Multimodal Learning Distributed Training DDP FSDP

✨ Generative AI

Diffusion Models Flow Matching SDXL FLUX Super-Resolution Flow Distillation One-Step Generation Image Generation Video Generation IP-Adapters Textual Inversion LoRA PEFT

πŸ‘οΈ Computer Vision

Object Detection Multi-Object Tracking Image Segmentation Video Segmentation Image Restoration Video Restoration Vision-Language Models Controllable Generation

πŸ€– Robotics And Robot Learning

MuJoCo NVIDIA Isaac Sim LeRobot

Vision-Language-Action Models Robotic Manipulation Visuomotor Policies Imitation Learning Reinforcement Learning Flow Matching Policies ManiSkill Robosuite Sim2Real

βš™οΈ Deployment And Infrastructure

NVIDIA Docker FastAPI Git Weights & Biases

NVIDIA Triton Inference Server MLflow GPU Optimization Model Serving Data Curation Experiment Tracking

πŸŽ“ Education

Bachelor of Science in Computer Science, University of Mumbai, 2022
CGPA: 9.21/10

🀝 Connect

I am interested in research and engineering conversations around generative vision, efficient deep learning, multimodal intelligence, and learning-based robotics.

Research depth. Engineering rigor. Systems that run outside the notebook.

Pinned Loading

  1. flow-based-models flow-based-models Public

    Conditional Flow matching for Image and video

    Python 13 1

  2. Zero-DCE Zero-DCE Public

    Implementation of Research Paper "Learning to Enhance Low-Light Image via Zero-Reference Deep Curve Estimation"

    Python 49 8

  3. PRANet-Polyps-Segmentation PRANet-Polyps-Segmentation Public

    Implementation of research paper : "PraNet: Parallel Reverse Attention Network for Polyp Segmentation" in Tensorflow

    Python 24 2

  4. mini-pi0 mini-pi0 Public

    Very minimal Implementation of Pi-Zero (https://www.pi.website/download/pi0.pdf). Vision Action Flow matching model

    Python 4