Aidan Erickson

Aidan Erickson

Machine Learning Research Engineer · New York, NY

Hi, I'm Aidan!

I am a Machine Learning Research Engineer at Wynd Labs, where I work on democratizing ML-enabled internet-scale datasets.

My education is in AI, computer science, and mathematics. I did my M.S at Carnegie Mellon where I studied Generative AI, Deep Learning, Speech Recognition, and Natural Language Processing.

I also contribute to open-source research at LAION. Most recently that's LAION-Tunes, an open music dataset and benchmark accepted to NeurIPS 2026.

In my free time, I like to attend astrophysics seminars at my local bar, and take astrophotography pictures of the stars and planets. My favorite planet is Saturn 🪐.

Thanks for visiting my site!

Learning-rate sweep: a 0.8M-parameter character-level GPT trained on Tiny Shakespeare

Train loss

  • Learning rate
  • 1e-4
  • 3e-4
  • 1e-3
  • 3e-3
  • 1e-2
1.52.02.53.03.54.0
00.5k1k1.5k

Training step

Research

Papers and research projects I've worked on.
  1. NeurIPS 2026

    • Audio
    • Datasets
    • Evaluation

    LAION-Tunes

    LAION-Tunes: An Open 1.4M-Track Dataset and Perceptual Benchmark for AI-Generated Music

    An open 1.4M-track dataset and a perceptual benchmark for AI-generated music.

    Robert Kaczmarczyk, Tawsif Ahmed, Felix Friedrich, Aidan C. Erickson, Orian Sharoni, Dorien Herremans, Christoph Schuhmann

    My part
    To add: your contribution
    Result
    To add: key findings
    • Paper: paper link
    • Dataset: dataset link
  2. Wynd Labs / Grass · 2025

    • Vision-language models
    • Fine-tuning

    Cliptagger

    A fine-tuned Gemma3-12B for video keyframe annotation.

    My part
    Trained the model.
    Result
    Beats Claude 4 Sonnet at video keyframe annotation at 1/20th the price.
    • Model: model link
  3. IEEE ITSC 2025

    • Autonomous vehicles
    • Perception

    Spectrum learning for low-cost perception

    Toward a Low-Cost Perception System in Autonomous Vehicles: A Spectrum Learning Approach

    A spectrum learning approach to low-cost perception for autonomous vehicles.

    Mohammed Alsakabi, Aidan C. Erickson, John M. Dolan, Ozan K. Tonguz

    My part
    To add: your contribution
    Result
    To add: key findings
    • Paper: paper link
  4. LAION · 2025 — Present

    • Audio
    • Contrastive learning
    • Open source

    Contrastive language-audio pre-training

    Open-source CLAP research at LAION.

    My part
    To add: your contribution
    Result
    To add: key findings
    • Code: code link

Timeline

Where I've worked and studied. Click an entry to see what I did there.
WorkEducation
  1. 2025–now

    • Designed highly optimized and scaled machine learning annotation pipelines, processing billions of textual, audio, and video data for leading frontier labs. Created high-throughput ML pipelines in only a few days for quick customer turnaround.
    • Trained Cliptagger, a fine tune of Gemma3-12b, beating Claude 4 Sonnet at video keyframe annotation at 1/20th the price.
    • Designed webpage metrics and ranking system for AI usefulness. Deployed across 60k high-signal websites in 59 countries.
    • Created multimodal retrieval systems from internet-scale webcrawls (billions of embeddings) with our Grass proxy network.
  2. 2025

    • Created demand forecasting and dynamic pricing models to optimize stay prices on over 6M reservations on 50k listings.
    • Designed, created, and stress tested system integrations for customers with limited documentation and 3-day deadlines.
    • Developed test environments to simulate high-load onboarding before making the platform publicly accessible.
  3. 2024

    • Master of Science: Artificial Intelligence. Graduated 12/2024.
    • GPA: 3.78/4.0
    • Coursework: Generative AI | Deep Learning | Speech Recognition | Natural language processing | Systems and Toolchains for AI
  4. 2024

    • Developed a convolutional graph neural network to detect identity fraud at scale (~2.6B Entries), tightening fraud analysis.
    • Conducted experiments and research on different methodologies, including transformer and convolutional recurrent models, among others, achieving an identity fraud prediction accuracy of over 89%. Used Neo4J and PyTorch Geometric on CUDA.
  5. 2023

    • Bachelor of Science: Computer Science, Minor in Mathematics. Graduated 5/2023.
  6. 2022

    • Implemented asset configuration features in GoLang, affecting 3,200 low orbit broadband satellites and ground equipment.

Volunteering

  1. 2025–now

    • Worked on multiple open-source research projects, including CLAP (Contrastive Language Audio Pre-Training) and Project Alexandria, a large-scale effort to free research from copyright into a unified, comprehensive knowledge base.
  2. 2023

    • Taught underrepresented, at-risk girls how to reason and program in Python for a summer. Increased test scores by 40%.

Projects

Things I built during my master's at Carnegie Mellon.

Carnegie Mellon · Fall 2024

Llama-2 Implementation

Implementation of the Llama-2 LLM in PyTorch.

  • PyTorch
  • Transformers
  • RoPE
  • GQA
  • Code: GitHub link

Carnegie Mellon · Fall 2024

Sheet Music Diffusion Generator

Diffusion model that generates sheet music.

  • Diffusion
  • U-Net
  • VAE
  • Code: GitHub link

Carnegie Mellon · Spring 2024

Sketch-to-Image Diffusion Network

Diffusion network that generates hand-drawn sketches from photographs.

  • Diffusion
  • Code: GitHub link

Carnegie Mellon · Fall 2024

Text-to-Image Diffusion Network

Text-to-image diffusion model with cross-attention.

  • Diffusion
  • Cross-attention
  • Code: GitHub link

Carnegie Mellon · Spring 2024

Host-Based Anomaly Intrusion Detection

Detects malware from Linux process syscall activity.

  • Security
  • AWS EC2
  • Code: GitHub link

Other: NumPy neural network · Radar denoising GAN · C2 prediction web app · Compiler

Contact

Send me a message here and I'll get back to you. I'm happy to talk about ML research, data pipelines, or anything on this page.