Classical Computer Vision + OCR · Public GitHub + Kaggle Project
Handwriting OCR
HOG Features · KNN Classification · EMNIST Letters
A classical OCR pipeline for handwritten English letter recognition using HOG feature extraction and a distance-weighted KNN classifier, extended with character segmentation and custom handwritten word reconstruction.

OVERVIEW
A classical OCR pipeline built around shape features
This project recognizes handwritten English letters using Histogram of Oriented Gradients (HOG) features and a K-Nearest Neighbors classifier. EMNIST Letters provides the official evaluation set, while an additional isolated-character dataset enriches training.
The notebook extends isolated-character recognition into a practical OCR workflow by segmenting custom handwritten word images into letters, classifying each character, and reconstructing the predicted word.
DATASET & FEATURES
RECOGNITION PIPELINE
Word image → character segmentation → HOG → KNN → text
Custom word images are separated into individual handwritten character regions.
Each normalized 28×28 letter is converted into a 1,296-dimensional gradient-orientation descriptor.
A distance-weighted KNN model with 7 neighbors predicts the corresponding English letter class.
Predicted characters are placed back in reading order to form the recognized word.
OFFICIAL EVALUATION
The official score is measured on the EMNIST Letters test split. The additional handwritten-character dataset is used to enrich training rather than replace the official test set.
PROJECT GALLERY
From benchmark letters to real handwritten words
The gallery combines outputs from the completed Kaggle notebook with a real handwriting demonstration sheet, showing both character-level evaluation and end-to-end word reconstruction.
Open full size ↗
Open full size ↗
Open full size ↗
Open full size ↗
Open full size ↗
Open full size ↗WHAT THIS PROJECT DEMONSTRATES
Feature engineering and classical ML remain useful for OCR
The project emphasizes a complete classical computer-vision workflow: image preparation, handcrafted HOG features, KNN classification, benchmark evaluation, segmentation logic, and visual inspection of errors on custom handwriting.
TECH STACK
PUBLIC PROJECT
Code, trained model release, and full notebook.
The GitHub repository contains the lightweight demo structure and links to the pre-trained KNN + HOG model release. Kaggle contains the complete training and evaluation notebook.