Skip to content
View Salik-Devv's full-sized avatar

Block or report Salik-Devv

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Salik-Devv/README.md

Mohammad Salik Dev

AI Hardware Acceleration · Computer Architecture · FPGA & GPU Systems

B.Tech, Electronics & Communication Engineering — Jamia Millia Islamia, New Delhi (2027)

Research Intern, Indian Institute of Technology Jodhpur

Email LinkedIn GitHub


About

My work focuses on the intersection of computer architecture, GPU/CUDA systems, and machine learning — building hardware accelerators that make LLM inference faster and more power-efficient. My recent work spans CUDA kernel optimization, FPGA-based GEMV accelerators using Vitis HLS, and RTL design, evaluated against a common benchmark: LLaMA-2 inference.

I'm particularly interested in AI hardware acceleration, hardware/software co-design, FPGA systems, and high-performance computing. My long-term goal is to contribute to R&D in computer architecture and AI systems.


Research Interests

  • FPGA Acceleration & Hardware/Software Co-Design
  • LLM Inference Optimization
  • Computer Architecture
  • Embedded AI
  • CUDA Programming & GPU Performance Engineering
  • High-Performance Computing
  • VLSI & Digital Design

Technical Skills

Languages
C C++ Python CUDA VHDL Verilog Assembly MATLAB

Hardware & FPGA
Vivado Vitis HLS Quartus ModelSim Questa Zynq

GPU & Parallel Computing
CUDA cuBLAS Nsight OpenMP

Tools
Linux Git CMake OpenCV NumPy Pandas


Research Experience

Research Intern — Indian Institute of Technology Jodhpur Hybrid FPGA/CPU acceleration for transformer inference — FPGA accelerator design in Vitis HLS, LLM inference optimization, power profiling, and roofline analysis.

Intern — NIT Meghalaya VLSI-focused internship.


Contact


"The best architectures are the ones that make the constraints of the silicon disappear into the performance of the algorithm."

Pinned Loading

  1. edge-detection-using-cuda edge-detection-using-cuda Public

    High-performance Sobel edge detection using CUDA with CPU vs GPU benchmarking, roofline analysis, and Nsight profiling.

    Python 4 3

  2. riscv-rv32i-pipelined-cpu riscv-rv32i-pipelined-cpu Public

    RV32I 5-stage pipelined CPU in VHDL with functional verification and core Fmax optimization

    VHDL 3

  3. llama2-cuda llama2-cuda Public

    LLaMA-2 CUDA C++ inference with GPU power profiling — RTX 4060 Laptop

    Cuda 2

  4. LLaMA-2-FPGA-Inference-Accelerator LLaMA-2-FPGA-Inference-Accelerator Public

    LLaMA-2 hybrid FPGA/CPU inference on Zynq UltraScale+ ZCU102, with PMBus power profiling.

    C++ 7 1