Applied AI · Public GitHub Release

Tetris AI Lab

End-to-End AI Gameplay Agent

AI agent trained from expert gameplay and strengthened with recovery training, Top-5 one-piece lookahead, and a limited safety filter.

RoleSole Developer
Year2026
StatusPublic GitHub Release
Tetris AI Lab AI gameplay agent project cover
PUBLIC RELEASE · INDIVIDUAL PROJECTCNN policy + Top-5 lookahead + limited safety filter

OVERVIEW

From learned placements to a complete AI system

Tetris AI Lab is an end-to-end machine-learning project that studies how an AI agent can learn placement behavior from expert gameplay and become more reliable in long-running play. The project began with a CNN placement policy trained from video-derived Classic Tetris examples and evolved into a hybrid agent combining learned candidate rankings with recovery-focused training, simulator-verified augmentation, one-piece lookahead, and a limited safety filter.

The final result is a complete Python/Pygame application with autonomous AI play, human play, Human vs Final AI, Greedy vs Final AI, replay viewing, screenshots, settings, and reproducible evaluation workflows.

PROJECT GOAL

More than predicting the next move

The objective was not simply to train a neural network to classify Tetris placements. The broader challenge was to build a testable AI system that could learn from expert demonstrations, survive imperfect states, reveal failure modes, and show whether each engineering change produced measurable improvement in actual gameplay.

Learn

Learn placement behavior from expert gameplay.

Recover

Expose the model to difficult states that expert demonstrations rarely contain.

Reason

Evaluate high-probability model suggestions with limited one-piece lookahead.

Measure

Evaluate the deployed agent using survival behavior in addition to offline model metrics.

DATA PIPELINE

From tournament footage to reproducible training data

The training workflow was built from video-derived Classic Tetris examples rather than a manually prepared toy dataset.

01Video-derived examples

Gameplay footage was converted into examples representing board states and expert placement decisions.

02Automated cleaning

Repeatable rules removed invalid or unsuitable examples before training.

03Duplicate control

Semantic duplicate handling reduced unwanted repetition in the training data.

04Video-level splitting

Train, validation, and test partitions were separated by source video to reduce leakage.

AI APPROACH

A learned policy strengthened with controlled reasoning

01

CNN Placement Policy

The neural network ranks candidate placements from the current board state and incoming piece information. It remains the primary learned component of the final agent.

02

Recovery Training

Simulator-generated difficult states expose the policy to situations that are uncommon in strong expert demonstrations but common after AI mistakes.

03

Verified Augmentation

Mirror augmentation is validated through the simulator so transformed examples remain consistent with legal Tetris behavior.

04

Top-5 Lookahead + Safety

The final agent evaluates the CNN's highest-ranked candidates using one-piece lookahead and applies a limited safety override to reduce catastrophic placements.

PERFORMANCE EVOLUTION

Each stage was evaluated as an agent, not only as a classifier

The benchmark summaries show a clear increase in average survival as recovery, safety, and lookahead components were added.

Tetris AI Lab agent performance evaluation chart
Random25.74
CNN Baseline44.45
Recovery v276.55
Recovery v3143.42
Safety Only307.32
Lookahead Top-3908.76
Lookahead Top-51,716.85
Final Agent1,846.60

Average pieces survived, as reported in the project benchmark summaries. The final comparison policies use a 100-episode / 2,000-piece benchmark; earlier stages retain their original evaluation settings.

WHY THE HYBRID AGENT MATTERS

Greedy predictions can look good locally and still fail over time

Small structural mistakes accumulate in a sequential environment. The enhanced agent keeps the CNN as the learned policy while allowing limited decision-time reasoning to reject some high-impact choices.

Greedy CNN compared with lookahead and safety Tetris agents

DEMO VIDEO

See the application and agent in motion

The demo shows the project interface, gameplay modes, and final AI behavior inside the released desktop application.

APPLICATION MODES

Built as an interactive application, not only a training notebook

AI Agent

Runs the final policy with live score, line, piece, timing, confidence, and safety information.

Human Play

Allows the same environment to be played directly by a human.

Human vs Final AI

Compares human and AI gameplay under the same seed and generator.

Greedy vs Final AI

Shows the practical difference between direct CNN selection and the enhanced agent.

Replay Viewer

Supports replaying recorded sessions for inspection and comparison.

Settings & Utilities

Includes runtime controls, screenshots, help, about information, and configurable gameplay settings.

PROJECT GALLERY

Interface, comparisons, and system behavior

ENGINEERING LESSONS

The gap between a good model and a good deployed agent

Distribution shift

Expert demonstrations mostly contain healthy boards. AI errors can move the game into unfamiliar states, motivating recovery-focused training.

Offline accuracy is not enough

A plausible single move can still damage long-term survival, so gameplay evaluation was treated as a separate measure of system quality.

Error accumulation

Small placement errors compound over hundreds of pieces, which made limited lookahead and safety checks useful.

Learned vs engineered behavior

The final design preserves the CNN as the main policy and uses additional logic selectively rather than replacing learning with a hand-built heuristic bot.

TECH STACK

Tools used across the full workflow

PythonPyTorchCNNPygameData Pipelines

ProgrammingPython

Machine LearningPyTorch · CNN

ApplicationPygame

Data & ExperimentationVideo-derived data · cleaning · splitting · recovery data · augmentation · benchmark tracking

Agent EngineeringTop-K ranking · one-piece lookahead · safety filtering

DevelopmentGit · GitHub · debugging · testing · documentation

NEXT STEPS

Potential directions for future experimentation

  • Deeper multi-piece planning
  • Learned value estimation
  • Expanded recovery-state generation
  • Larger-scale gameplay evaluation
  • Additional ablation studies
  • Further analysis of offline metrics vs long-horizon gameplay

PUBLIC GITHUB RELEASE

Explore the source, release notes, and project documentation.