A high-performance, full-stack AI microservice demonstrating low-level GPU acceleration. CUDA-NN-App bridges custom C++/CUDA kernels directly into PyTorch via C++ extensions, serving low-latency RESTful inference endpoints with FastAPI.
- Custom CUDA Kernel: Fused Linear matrix multiplication, bias addition, and ReLU activation executed in a single GPU pass to reduce memory bandwidth bottlenecks.
- PyBind11 Integration: Seamless C++-to-Python bindings compiled Just-In-Time (JIT) using PyTorch's
torch.utils.cpp_extension. - Async REST API: Asynchronous inference service powered by FastAPI and Uvicorn.
- Production Layout: Clean separation of GPU compute, C++ bindings, network architecture, and service handlers.
cuda-nn-app/
├── cuda_ops/
│ ├── fused_linear_relu.cu # CUDA Kernel (Matrix Mult + Bias + ReLU)
│ └── bindings.cpp # C++ Host function & PyBind11 bindings
├── app/
│ ├── model.py # PyTorch Module integrating custom CUDA op
│ └── main.py # FastAPI application & REST endpoints
├── .gitignore # Git ignore rules
├── requirements.txt # Dependencies
└── run.py # Application launcher script
Ensure your host system meets the following prerequisites:
- OS: Linux / Windows with WSL2
- GPU: NVIDIA GPU (Compute Capability 6.0+)
- NVIDIA Drivers & CUDA Toolkit: CUDA 11.8 or higher installed with
nvccaccessible in PATH. - C++ Compiler:
g++(Linux) orMSVC(Windows) supporting C++17. - Python:
3.10or higher.
git clone https://github.com/YOUR_USERNAME/cuda-nn-app.git
cd cuda_nn_appCreate and activate a virtual environment inside the project root:
python3 -m venv .venv
# On Linux / macOS:
source .venv/bin/activate
# On Windows (PowerShell):
\.venv\Scripts\Activate.ps1Install ninja (for fast JIT compilation), PyTorch (with CUDA support), and FastAPI:
pip install --upgrade pip
pip install -r requirements.txtNote: Ensure the installed PyTorch CUDA version matches your system's CUDA Toolkit version (
python -c "import torch; print(torch.version.cuda)").
Start the API server using the launcher script:
python run.pyThe first startup will automatically invoke JIT compilation of the C++/CUDA code via torch.utils.cpp_extension. Subsequent starts will use cached binaries.
When the server is running, interactive API documentation is available at:
- Swagger UI:
http://localhost:8000/docs - ReDoc:
http://localhost:8000/redoc
Verifies GPU accessibility and CUDA availability.
Response:
{
"status": "active",
"cuda_available": true,
"device_name": "NVIDIA GeForce RTX 3080"
}Executes forward pass inference through the custom CUDA-backed neural network.
Request Body:
{
"data": [
[0.12, -0.43, 0.88, ...], // 128-float input vector
[-0.05, 0.91, -0.12, ...]
]
}Example Request (curl):
curl -X 'POST' \
'http://localhost:8000/predict' \
-H 'Content-Type: application/json' \
-d '{
"data": ['"$(python3 -c "import json, numpy as np; print(json.dumps(np.random.randn(1, 128).tolist()))")"']
}'This project is open source and available under the MIT License.