Conversion Optimization
Published · 11 min read
In 1966, two psychologists named David Green and John Swets published a book that grew out of wartime radar research. The problem they formalized was simple and brutal. A radar operator stares at a blip. A real aircraft and a random noise spike produce the same blip. Nothing in the signal itself tells the operator which one it is. Whatever separates them has to come from somewhere other than the screen.
Session recordings don't show why users drop off. They were never built to. A replay is a faithful and complete record of one thing, and the question a growth team brings to it is about something else entirely, which is why watching more of them feels like progress and so rarely is.
Here is the transfer. A session recording is a sensor, and what it captures is motion: cursor paths, scroll depth, field focus, the pause before a click. Intent was never in that signal. Two visitors in opposite internal states, one confused and one carefully reading, produce very nearly the same trace. The operator watching replays is in Green and Swets' position exactly. The blip is real. The interpretation is being supplied by the person watching, not by the sensor, and that person has no way to check their own hit rate.
Find out why your page loses people
Drop a URL. Get the first audit free when Nudgent opens.
This is not a criticism of Hotjar. Hotjar drop off analysis is genuinely good at the thing it was built for. Scroll maps, click maps, rage-click flags, funnel step exits. If you want to know that 61% of visitors who land on step two never reach step three, a recording tool will tell you, and it will tell you fast, and it will be right.
The trouble starts one question later.
You watch the recordings. You see forty people reach the company-size dropdown and stop. Some of them scroll back up. A couple click into the field and click out. One sits for eleven seconds with no input at all. You now know the drop-off is at the dropdown with a precision no survey would give you.
What you do not know is which of these is true:
Every one of those produces roughly the same replay. That is the point. It is not that the tool is missing a feature. It is that intent is not present in motion data, so no amount of replay volume recovers it. This is a category limit, in the way that a thermometer has a category limit about humidity.
The repo's community research turned up 12 independent versions of this same complaint across G2, Reddit, Hacker News and analytics blogs. The line that stuck was somebody describing having to watch hundreds of recordings to find one relevant answer. Not "the tool is bad." Just the quiet cost of using a location instrument to do inference work.
When the sensor doesn't carry the signal, teams bridge the gap manually. There are three standard bridges, and each one has an honest cost that rarely gets said out loud.
Bridge one: watch more recordings. The intuition is that volume will eventually surface the pattern. Sometimes it does, for the narrow class of problems that are visible as motion: a button that doesn't respond, a layout that breaks on a particular viewport, a validation error that fires silently. For everything else, volume mostly increases confidence without increasing accuracy. You watch sixty instead of forty, you see the same stall sixty times, and you walk away more certain of a conclusion you formed on recording number three. In signal detection terms, you've lowered your threshold for calling a target. Your hit rate goes up. So does your false-alarm rate, and nothing in the stack is counting those.
Bridge two: talk to users. This one genuinely works. A usability session with five people will teach you more about why users abandon signup than five hundred replays. The cost is time and sample. Recruiting, scheduling, running, and synthesizing takes a week or two if you are quick about it, and five to eight participants is not your funnel. It is a sample that is excellent for finding whether a problem exists and poor at telling you which of four known problems is costing the most. Most teams run this once per quarter at best, which means it is not available for the decision you are making on a Tuesday afternoon.
Bridge three: apply best practice. Shorten the form. Add logos. Move the CTA above the fold. These are fine defaults, and they are fine precisely because they are averages taken across many pages that are not yours. They tell you what usually helps. They cannot tell you whether the thing usually helping is the thing currently broken here. We have seen pages where cutting form fields did nothing because the fields were never the constraint, and the actual problem was that the page never established who the product was for, so nobody reading it could tell whether they qualified.
All three bridges share a structure. They are attempts to recover intent from a system that never recorded it, by adding a second source of information. Interviews add a second source and work, slowly. The other two mostly add confidence.
A conversion health audit is the second instrument. It does not watch visitors at all. It measures the page.
That distinction matters more than it sounds. Recordings sample behavior and ask you to infer the page from it. An audit inspects the page directly and asks what a reasonable visitor would struggle with, given what is actually on the screen. Same starting artifact, a URL, entirely different measurement target.
Concretely, an audit reads the signup page and scores it against named conversion friction dimensions. Is the next action unambiguous. Is the value stated before the ask. Are the anxieties a person would reasonably have about handing over a work email addressed anywhere on the page. Can a visitor in your target segment see themselves described. Each of those is checkable without a single visitor, because each is a property of the page rather than a property of the traffic. Our audit methodology writes out how the scoring works, and the dimensions breakdown covers all 14, seven for humans and seven for the AI agents that increasingly read your pricing page on a buyer's behalf before the buyer ever sees it.
The output is not "61% exit here." It is closer to: decision clarity scores 4 because the form asks for company size before the page has established which sizes the product serves, and motivation strength scores 5 because the value proposition appears only after the form. Those are claims you can check against the page yourself in about ninety seconds, which is the honest test of whether an audit is telling you something or telling you a story.
Here is where I should be straight about the weakness in Nudgent's own method. An audit measures the page, not your audience. It cannot know that your buyers are procurement-heavy and will tolerate a longer form than the general benchmark suggests. It reasons from what is visible, which means it can flag something as friction that your particular segment handles fine. Recordings have the opposite failure: they see your actual audience and cannot explain them. Neither instrument is complete. They fail in different directions, which is the only reason running both is worth anything.
If you want to see the split drawn out against a specific tool, we've done it in the Microsoft Clarity comparison. The short version applies to Hotjar too. Keep the recording tool. It confirms where, and it will confirm whether your fix worked, which an audit cannot do.
You almost certainly already have the input.
If Hotjar has pointed you at a specific step, that URL is the thing to run an audit against. Not the homepage, not the funnel in aggregate. The one page where the recordings show the stall. The recordings did the localization work. The audit does the diagnosis work on the localized spot.
Then the sequence is boring and it is the whole point:
Step four is the part teams skip, and it is the only step that generates the data you actually lack. Every time you predict a cause, ship a fix, and measure the result, you learn something about your own accuracy. Do that six times and you have a rough sense of how often your reading of a recording was right. Do it zero times and you have six years of confident interpretation with no error rate attached, which is the state most growth teams are in, including ones with very good instincts. More on how we think about all of this in the glossary and across the blog.
Which brings me to the question I would rather not answer about my own pages.
The last five times you decided what was causing a drop-off, and acted on it, how many times were you right? Not "did the number move" but "did it move for the reason you said it would." And if you cannot answer that, is there anything currently in your stack that would have told you?
Hotjar tells you where, with real precision. It records cursor movement, scroll depth, clicks, rage clicks and exact funnel exit steps. It does not capture intent, because intent is not present in motion data. Two visitors in opposite mental states, one confused and one deliberately reading, produce almost identical recordings. The interpretation of why is supplied by whoever watches the replay, not by the tool, which means it carries that person's error rate rather than the sensor's.
Yes, and that is the intended pairing. Recordings localize the problem to a specific step. A conversion health audit inspects that specific page and explains what about it is likely causing the exit, ranked by likely impact. Then recordings verify whether the change moved the exit rate. Recordings answer where and whether it worked. An audit answers why and what to fix first. Neither one covers the other's blind spot.
A heatmap aggregates visitor behavior into a visual summary of attention and clicks. A conversion audit inspects the page itself and scores it against named dimensions such as decision clarity, trust signal and motivation strength. The heatmap needs traffic and shows you what happened. The audit needs only a URL and shows you what a reasonable visitor would struggle with. This also means an audit works on a page that has no traffic yet, which is where heatmaps have nothing to say at all.
Watching more recordings raises your confidence faster than it raises your accuracy. Recordings are reliable for problems that are visible as motion: broken buttons, layout failures, silent validation errors. For anything involving comprehension, trust or motivation, additional volume mostly confirms the location you already had. If you have watched twenty and you still cannot say why, the remaining eighty will not tell you either. That is a limit of the instrument rather than a limit of your effort.
Find out why your page loses people
Drop a URL. Get the first audit free when Nudgent opens.
Get early access