mahfujmustafa.dev / projects

Psychedelic Model

ML and neuroscience research, June to July 2026. Code on GitHub.

DMT research describes the visuals in a fairly consistent order. Geometry comes first, then forms, then figures, then a return to the room. I wanted to know whether that order falls out of an image model when the only things I touch are parameters that map to known brain variables.

So I started Stable Diffusion 1.5 from an ordinary photo and wrote five knobs that each stand in for something the brain does, then moved them along a dose curve. Nothing is painted in, the prompt never asks for anything psychedelic, and the metrics and their thresholds were written down before the runs they judge. It's a computational model of reported visual experience, so it can show a signature in the images but can't tell you what a trip feels like.

The idea

Two ideas from neuroscience set this up. REBUS (Carhart-Harris and Friston) says psychedelics relax the precision of high-level priors, so the brain stops explaining away bottom-up signal and internal noise. The other is 5-HT2A gain. DMT raises the gain of cortical pyramidal cells, which destabilizes a layered visual system. A latent diffusion model is loosely the same kind of system, with a prior (the text conditioning), sensory evidence (the encoded photo) and a stack of levels (the U-Net).

The five knobs

Everything comes from one dose curve that peaks at 4 minutes and never reaches zero.

The caption is written by BLIP, and then a filter strips 78 words like fractal, spiral and entity.

How a frame gets made

The VAE encodes the photo and the previous frame, and they get blended by scene_precision and recurrence. Noise is added, DDIM denoises it under guidance at prior_precision, and a forward hook multiplies the chosen U-Net block by gain. The result is decoded, contrast-matched to the photo, and becomes the previous frame for the next step.

What happened

The main run (40 frames, seed 0) reproduced the timing of the arc and none of the content of its phases. The image left the photo and came back, with RMSE against frame 0 peaking at 0.274 at frame 13 and back down to 0.025 by frame 39. Periodic patterns grew out of the noise at the peak, with excess power at a 19 px wavelength reaching 3.0x the natural spectrum against 1.6x for the original photo. But mirror and rotational symmetry stayed near zero (-0.015 to 0.049), while synthetic test patterns scored 0.98 to 0.99. That's periodicity without symmetry, so I don't call it a Klüver form constant. The figures never emerged.

Why parts of it failed

Raising the denoise strength drowns fine detail first, so the come-up blurred the photo instead of making it shimmer. In the brain the prior goes up and the evidence is left alone, while this engine raises the prior by wrecking the evidence.

The U-Net's normalization cancels a plain multiply, so what's left of gain is the tanh clip, and that makes blobs at whatever scale the block already works at. That block has one cell per 16 px, which is where the 19 px wavelength comes from. In the brain a lateral interaction kernel picks the wavelength, and a scalar knob has nothing like it.

The missing figures say nothing about DMT. Stable Diffusion's only route to a figure is the prompt, and I strip figure words out of it on purpose. The return is guaranteed by construction, since the photo is re-encoded clean every frame. And the headline numbers come from one run, on one photo, with one seed.

What's next

The development plan sets a pass or fail line for each experiment in advance. The first few are a come-up driven by gain alone with denoise strength held flat, jittering the input photo every frame to see if the return survives, and a difference-of-Gaussians lateral kernel to see whether a selected wavelength is reachable.

Stack

Python and PyTorch, with Stable Diffusion 1.5 and BLIP from Hugging Face, running on Apple Silicon. I used agentic AI tools for assisted coding and research and disclose that in the paper I'm writing from this project. The question, the design and the thresholds were mine.

Code and full write-up on GitHub, or back to mahfujmustafa.dev and the rest of my projects.