Machine learning research · London

Felix Baastad Berg

I work on why neural networks learn what they learn — the geometry of internal representations, when generalization actually arrives, and how generative models move noise into data. Currently an Aker Scholar finishing an MSc in Advanced Computing at Imperial College London.

Felix Baastad Berg
Imperial College London
2025 — 2026

MSc Advanced Computing as an Aker Scholar. Working on representation geometry in grokking and on generative models, supervised by Tolga Birdal.

MIT & Harvard
2024 — 2025

Fulbright Scholar in Mathematics at MIT, and a year with the Rajan Lab at Harvard building deep RL simulations to model animal behaviour.

NTNU
2019 — 2025

Five-year integrated MSc in Mathematics. Thesis on continuous attractor networks and memory in deep reinforcement learning.

Mathematics

Persistent homology, intrinsic dimension, and the manifolds representations live on.

Machine learning

Generative models, generalization, and deep reinforcement learning with memory.

Neuroscience

Grid cells, continuous attractors, and spiking models of biological memory.

Computing

JAX on GPU clusters, large-scale simulation, and numerical methods in C++.

01

Research

8 selected

01 Under review 2026 · NeurReps · Imperial College London

Auditing Geometric Change in Grokking: Centroid Configuration versus Within-Class Geometry

Geometric statistics move sharply when networks grok, but a global measurement cannot say what moved. Writing each activation as a class centroid plus a within-class residual and intervening on the components separately shows that MST-based PH-dimension is driven largely by the centroid configuration, while TwoNN goes invariant once its neighbours fall inside classes — so the two estimators can assign opposite meanings to the same representational transition.

persistent homologyintrinsic dimensiongrokkingneural collapse

02 Conference talk 2026 · Individual Study Option · Imperial College London

Predicting Grokking Onset from Intrinsic Training Metrics

Can you tell a network is about to generalize without ever looking at test error? Tracking intrinsic signals of the internal representations — correlation dimension, PCA participation ratio, persistent homology and MST-based fractal dimension — first-layer MST dimension leads the generalization gap by roughly 1000 epochs. A parametric onset predictor recovers grokking time from those dynamics alone (R² ≈ 0.98). Selected to present at the Imperial Department of Computing ISO Conference.

topologylead–lag analysisPyTorch

03 In progress 2026 · MSc thesis · Imperial College London

Transition Flows for Generative Modelling

Ongoing thesis work combining TarFlow — an autoregressive normalising flow over image patches — with Transition Matching. Instead of predicting a single deterministic direction the way flow matching does, the flow models and samples from the full distribution of plausible transition directions, keeping exact likelihoods while generating inside a single model.

normalising flowsflow matchingdiffusiontransformers

04 Published 2025 · NeurIPS · Harvard University

Deep RL Needs Deep Behavior Analysis

Two agents can earn the same return while solving the task in completely different ways, so a reward curve says almost nothing about what a policy has learned. We show that model-free agents display implicit planning in open-ended environments, and introduce a behavior-analysis toolkit that reveals how policies reason and generalize.

JAXreinforcement learningHPCstatistics

05 Thesis 2025 · MSc Mathematics · NTNU

Spatial Representation Learned in the Recurrent Memory of Artificial Neural Agents

Grid-cell activity lives on a torus: two periodic phases that a moving agent has to integrate. I built biologically inspired RL agents with LSTM memory and continuous attractor networks, designed analytical and numerical grid-cell modules in JAX, and analysed statistically how that toroidal structure appears in recurrent memory.

JAXattractor networksgrid cellsFourier basis

06 Project 2024 · MIT

Rethought Generalization: Empirical Analysis of Bounds for Compositionally Sparse Networks

An empirical study of norm-based generalization bounds in CNNs under varying data sizes, random labels and optimization strategies — showing when tighter bounds actually say something useful, and how mini-batch size and regularization move the picture.

generalization theoryCNNsPyTorch Lightning

07 Project 2025 · MIT

Unsupervised Time Series Forecasting with Spiking Neural Networks leveraging STDP

Spiking neural networks trained with STDP for time series forecasting, benchmarked against autoregressive models to evaluate predictive capacity on unclustered datasets.

spiking networksSTDPscikit-learn

08 Project 2024 · Harvard University

Deep Reinforcement Learning for Memory-Driven Navigation in Predator-Prey Environments

A custom predator-prey grid world for evaluating memory in RL agents. LSTM-based agents outperformed feedforward models, with behavioural analysis of resource revisiting and predator evasion.

LSTMreinforcement learningPyTorch
02

Experience

Co-Founder (CTO → COO)

· Kateter 2021 — present

A digital learning platform for scientific courses with roughly 10,000 users. Led development of an interactive math-visualization library on top of Three.js.

Machine Learning Intern

· Firda Jan — Aug 2023

LLM intern at a venture capital firm scaling Norwegian technology companies. Built a LangChain-based chatbot that returned verified sources to support investment and organizational decisions.

Data Scientist

· Heimstaden Jun — Dec 2022

Built predictive rent-price models in Python at Europe's second-largest residential real estate company.

Data Science Intern

· Atlas Jun — Aug 2021

Built functionality and algorithms for analysing geospatial satellite data with Django, GeoPandas and SQL, supporting renewable energy projects.

Simulation Lead

· Propulse NTNU Sep 2020 — Jun 2021

Led development of the flight simulator for a student-built rocket that placed 2nd of 75 teams at the Spaceport America Cup 30K COTS. C++ and computational fluid dynamics.

03

Honors

2025

Aker Scholarship

Norway's most prestigious graduate scholarship, fully funding advanced studies at leading global universities.

2024

Fulbright Scholarship

One of six candidates selected for the non-degree Fulbright fellowship, spent at MIT.

2019

Norwegian Physics Olympiad — National Finals

Reached the national finals, ranking among Norway's top high school physics students.