Skip to content

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

CUDA-NN-App

Python C++ CUDA PyTorch FastAPI

A high-performance, full-stack AI microservice demonstrating low-level GPU acceleration. CUDA-NN-App bridges custom C++/CUDA kernels directly into PyTorch via C++ extensions, serving low-latency RESTful inference endpoints with FastAPI.


Key Features

  • Custom CUDA Kernel: Fused Linear matrix multiplication, bias addition, and ReLU activation executed in a single GPU pass to reduce memory bandwidth bottlenecks.
  • PyBind11 Integration: Seamless C++-to-Python bindings compiled Just-In-Time (JIT) using PyTorch's torch.utils.cpp_extension.
  • Async REST API: Asynchronous inference service powered by FastAPI and Uvicorn.
  • Production Layout: Clean separation of GPU compute, C++ bindings, network architecture, and service handlers.

Project Architecture

cuda-nn-app/
├── cuda_ops/
│   ├── fused_linear_relu.cu    # CUDA Kernel (Matrix Mult + Bias + ReLU)
│   └── bindings.cpp            # C++ Host function & PyBind11 bindings
├── app/
│   ├── model.py                # PyTorch Module integrating custom CUDA op
│   └── main.py                 # FastAPI application & REST endpoints
├── .gitignore                  # Git ignore rules
├── requirements.txt            # Dependencies
└── run.py                      # Application launcher script

Requirements

Ensure your host system meets the following prerequisites:

  • OS: Linux / Windows with WSL2
  • GPU: NVIDIA GPU (Compute Capability 6.0+)
  • NVIDIA Drivers & CUDA Toolkit: CUDA 11.8 or higher installed with nvcc accessible in PATH.
  • C++ Compiler: g++ (Linux) or MSVC (Windows) supporting C++17.
  • Python: 3.10 or higher.

Quickstart Guide

1. Clone the Repository

git clone https://github.com/YOUR_USERNAME/cuda-nn-app.git
cd cuda_nn_app

2. Set Up Virtual Environment

Create and activate a virtual environment inside the project root:

python3 -m venv .venv

# On Linux / macOS:
source .venv/bin/activate

# On Windows (PowerShell):
\.venv\Scripts\Activate.ps1

3. Install Dependencies

Install ninja (for fast JIT compilation), PyTorch (with CUDA support), and FastAPI:

pip install --upgrade pip
pip install -r requirements.txt

Note: Ensure the installed PyTorch CUDA version matches your system's CUDA Toolkit version (python -c "import torch; print(torch.version.cuda)").

4. Run the API Server

Start the API server using the launcher script:

python run.py

The first startup will automatically invoke JIT compilation of the C++/CUDA code via torch.utils.cpp_extension. Subsequent starts will use cached binaries.


API Documentation

When the server is running, interactive API documentation is available at:

  • Swagger UI: http://localhost:8000/docs
  • ReDoc: http://localhost:8000/redoc

Endpoints Overview

1. Health Check (GET /health)

Verifies GPU accessibility and CUDA availability.

Response:

{
  "status": "active",
  "cuda_available": true,
  "device_name": "NVIDIA GeForce RTX 3080"
}

2. Model Inference (POST /predict)

Executes forward pass inference through the custom CUDA-backed neural network.

Request Body:

{
  "data": [
    [0.12, -0.43, 0.88, ...],  // 128-float input vector
    [-0.05, 0.91, -0.12, ...]
  ]
}

Example Request (curl):

curl -X 'POST' \
  'http://localhost:8000/predict' \
  -H 'Content-Type: application/json' \
  -d '{
    "data": ['"$(python3 -c "import json, numpy as np; print(json.dumps(np.random.randn(1, 128).tolist()))")"']
  }'

License

This project is open source and available under the MIT License.

About

High-performance neural network inference engine powered by custom C++/CUDA fused kernels and served via FastAPI REST APIs.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages