Search for this and you will find confident numbers. One app’s marketing page says webcam tracking lands “within 2-5 centimeters” of marker-based motion capture and credits “Google Research (2023)” without linking to anything. Another says 98%. A third says 95.8%. None of them agree, and almost none of them say what was measured, on whom, or in what room.
The honest answer takes a bit longer, and it starts with an awkward admission: measured against a proper motion capture lab, webcam pose estimation is not very accurate at all. A study published in November 2025 put the gap at 7 to 12 centimetres per joint. That is roughly the width of your hand.
That number is real, and it is also the wrong thing to worry about. What follows is what the peer-reviewed research actually found, why lab accuracy is a poor guide to whether an app will nag you correctly at 3pm on a Tuesday, and a ten-minute test that tells you more about your own setup than any published figure will.
The Short Answer
Three separate things get called “accuracy” here, and they behave nothing alike.
Joint position accuracy asks how close the model’s guess at your shoulder is to where your shoulder actually is. Against optical motion capture it is poor. Errors of several centimetres are normal.
Joint angle accuracy asks how close it gets to the real angle at a joint. Better. Still short of clinical standards, and it swings wildly depending on which joint you pick and where the camera sits.
Posture classification accuracy asks whether the app correctly decides you are slouching. With your upper body clearly in shot, that runs into the nineties in study conditions. For anything the desk hides, it collapses.
Your posture app only ever needed the third one, and it needed it for you rather than for people in general. Almost every confusing accuracy claim you will read comes from someone quoting the first two at you.
What Lab Validation Studies Found
The most thorough recent test is Assessment of monocular human pose estimation models for clinical movement analysis, published in Scientific Reports in November 2025 by a team from ETH Zurich, Akina AG and the Schulthess Klinik. They built a dataset called Physio2.2M: 2.2 million video frames of 25 people doing physical exercises, each frame paired with ground truth from a passive-marker optical motion capture system. Then they ran 18 different 2D pose estimators and 26 different 3D ones against it.
The headline figures are sobering. Mean per joint position error came out “in the range of 72 to 122 mm in 2D within the image plane and 146 to 249 mm in 3D when considering depth.” For knee flexion angle, the best 2D model managed a mean absolute error of 9.3 degrees and the rest were worse. The authors are blunt about what that means clinically: joint angle errors need to be small enough for correct clinical interpretation, “which none of the pose estimators tested was able to achieve.”
Detection ratios in that same paper ranged from 56.50% to 100%. Some of these models simply failed to find a person at all in nearly half the frames they were shown. Depth is the other sore point. The 3D error runs at roughly double the 2D error, which is what you would expect from one camera trying to guess distance.
Now the counterweight. A 2022 study in Gait & Posture, Comparing the accuracy of open-source pose estimation methods for measuring gait kinematics, tested OpenPose, MoveNet Lightning, MoveNet Thunder and DeepLabCut on 32 walking participants against marker-based capture. OpenPose and MoveNet Thunder were the most accurate for hip kinematics, averaging 3.7 ± 1.3 and 4.6 ± 1.8 degrees of error respectively across the whole gait cycle. OpenPose managed 5.1 ± 2.5 degrees on knee kinematics.
So: 4.6 degrees of error on hip flexion in one study, 9.3 degrees on knee flexion in another, from the same broad family of technology. What changed was the task, the camera placement, the plane of movement and which joint got asked about. Anyone quoting a single accuracy figure for “webcam posture detection” without those details is quoting a number that means very little.

Why the Lab Number Is the Wrong Yardstick
A motion capture lab wants to know where your knee is, in millimetres, in a room. Your posture app wants to know whether you are more slumped now than you were twenty minutes ago. One of those is a much easier question, and it is the one you are paying for.
Consider what stays fixed at your desk. The camera does not move. Lighting barely shifts across a working session, you are in the same jumper you had on an hour ago, and the chair has not gone anywhere. Under those conditions a systematic error in where the model thinks your shoulder sits mostly cancels itself out, because the app is comparing your shoulder now against your shoulder at calibration and both readings carry the same bias. The error is still there. It just stops mattering.
This is why the calibration step exists. A well-built app captures your upright and your slouch, builds a small classifier from those examples, and from then on measures deviation from your own baseline rather than from an idealised spine. The mechanics of how that pipeline works are worth understanding if you want to judge an app properly, because the pose model is only the first stage of it.
The trade-off is that this only holds while the setup holds. Move the laptop to the kitchen table and the calibration is worthless. Swap a fitted shirt for a heavy hoodie and shoulder keypoints shift. Sit sideways to take a call and the geometry the app relies on stops describing you.
Which makes the useful question a different one entirely. How stable is this within a single sitting, and how gracefully does it fail when something moves?

What Seated Posture Studies Report
Studies that test the actual job give more useful numbers. What the camera can see clearly, it classifies well. Everything else falls apart.
Kapoor, Jaiswal and Makedon presented a light-weight seated posture guidance system at PETRA 2022. Their pipeline takes frames from an ordinary webcam or phone camera, extracts body landmarks, and runs a neural network classifier over all 33 keypoints. It distinguishes good from bad seated posture with 98% accuracy. That is the sort of figure that ends up in marketing copy, and it is a real result from a real paper.
Compare it with PoseTrack, described in an August 2025 arXiv paper on a MediaPipe-based sitting posture monitor. The researchers recorded four seated postures from four camera perspectives and reported accuracy per posture rather than as a single average:
| Posture | Detection accuracy |
|---|---|
| Leaning forward | 100% |
| Good posture | 70% |
| Leg on chair | 50% |
| Crossed legs | 0% |
Crossed legs failed in every single video. The authors explain why: “the knees often being blocked by the table in the diagonal perspective and the arms in the left/right perspective,” with the result that “Mediapipe is often unable to correctly label the positions of the user’s knees.”
Forward lean, which is the single most common desk slouch, is the thing these systems detect best. Anything below the desk edge is close to invisible to them. So an app that claims to track your leg crossing from a webcam mounted on your monitor is claiming something the published testing says it cannot do.
Five Things That Move Accuracy in Your Actual Room
Published figures come from controlled conditions. Yours will not be. Five variables do most of the damage.
Camera angle. Pose models are trained largely on people photographed from a fairly natural viewpoint. A webcam clamped high above you, tucked low beside the keyboard, or shoved off to one side hands the model a perspective it has seen far less of. Off-axis views cause a second, sneakier problem. They compress the very distances a posture app measures, so when you are turned side-on the horizontal gap between your eyes shrinks towards nothing and any ratio built on that measurement goes haywire. This is the number one cause of an app nagging you while you are sitting perfectly well.
The desk itself. Everything from roughly your ribs down is behind a wooden panel. Hip keypoints get guessed rather than seen, and the PoseTrack results show what happens next. Judge an app on what it claims about your head, neck and shoulders, and discount anything it says about your pelvis.
Image quality, but less than you think. A 2022 study in the Journal of Imaging on the effects of image quality on human pose estimation accuracy systematically degraded resolution, colour depth and noise. Pose estimation held up better than most people expect. Accuracy stayed acceptable down to human subjects occupying roughly 60 × 80 pixels, and did not fall off badly until colour depth dropped below 512 colours. Noise was the harsher one. Error started climbing once 5% of pixels had been replaced with noise. Sensor noise is exactly what a cheap webcam produces in a dim room, so dim lighting hurts you through grain rather than through darkness.
What you are wearing and what is behind you. Loose or bulky clothing obscures where the joint actually sits. A busy background, a bookshelf with vertical lines, or another person walking through the frame all give the model competing candidates. The single-pose MoveNet variants most apps use track one person at a time, so a housemate crossing behind you is a genuine failure mode rather than a theoretical one.
Who you are. This one gets skipped and it should not. Sony AI’s write-up on limitations in fairness evaluations for human pose estimation evaluated OpenPose, MoveNet and PoseNet and found the underlying evaluation data badly skewed. Lighter-skinned individuals appeared over ten times as often as darker-skinned ones, and darker-skinned women appeared in only 17 of 657 valid images, around 2.6%. Automated and manual demographic annotations diverged by up to 70% for skin tone, which means even the fairness measurements are shaky. The honest position is that per-group accuracy for these models is not well characterised, and if an app performs poorly for you where it works for a colleague, the model is a plausible culprit rather than your imagination.
Be Sceptical of Any Single Percentage
Go back to that “within 2-5 centimeters of marker-based systems” claim. Set it beside the Scientific Reports figure of 72 to 122 mm, which is 7.2 to 12.2 centimetres, from 2.2 million frames with published methodology. The marketing claim is off by a factor of two to five, attributed to research that is never linked.
A few questions cut through most of this.
- Measured against what? Marker-based optical motion capture is the only ground truth worth quoting.
- On how many people, and were they all shaped and lit like the people in your office?
- Which posture? An average across four postures conceals a 0% and a 100%.
- Which joints? Shoulders and ears are reliable. Hips behind a desk are not.
- Is it a classification score or a position error? They are different units and they get muddled constantly.
An app that publishes none of this is not necessarily bad. It just has not given you anything to check, which puts you back on testing it yourself.

Test Your Own Setup in Ten Minutes
Your room, your camera, your body. Run this on any app before you decide whether it works.
- Sit normally and check the camera can see you. Head and both shoulders in frame, camera roughly at eye level, nothing across the lens. If your monitor height is wrong the camera angle usually is too, and fixing the screen fixes both.
- Calibrate in your real working position. Not your best posture. The way you sit when you are concentrating, in the clothes you actually wear, with the lighting you actually have.
- Run five deliberate slouches. Drop your head and round your shoulders, hold it, and count how long the app takes to notice. Five out of five should be caught. Four is workable. Three means recalibrate.
- Then sit well for thirty minutes and count the false alarms. This is the test that matters, and it is the one people skip. A nudge that fires while you are sitting fine is far more damaging than a slouch that goes unnoticed, because it teaches you to ignore the app.
- Apply the one-in-five rule. If more than roughly one nudge in five arrives while you are sitting properly, something is wrong with the setup rather than with you. Check camera angle first, lighting second, then recalibrate.
- Repeat once in the afternoon. Light changes. A setup that works at 9am in winter can behave differently at 4pm with a window behind you.
That takes ten minutes and it measures the only configuration you will ever actually use, which is more than any published accuracy figure can claim.
What “Accurate Enough” Actually Means
There is a version of this technology that is extremely accurate and completely useless. It sits in a lab with eight cameras and costs more than your car.
For a desk app the bar is much lower and shaped differently. Catching the slouch you actually do, which is overwhelmingly a forward lean with a dropped head, is the easiest thing on the whole list. Staying quiet while you are sitting fine is harder. Hardest of all, and the thing almost nobody advertises, is knowing when to say nothing at all, so that when you swing side-on to talk to someone the app holds its fire rather than firing wrongly.
A model that recognises it cannot see you properly, and shuts up, beats a marginally more accurate one that guesses. Whether a posture app is worth it usually comes down to whether you still believe its nudges in week three, and false positives are what destroy that belief.
SitApp uses Google’s MoveNet, which the TensorFlow team benchmarks as running “faster than real time (30+ FPS) on most modern desktops, laptops, and phones,” and runs it entirely on your own machine so no frames ever leave the device. On top of that sits a classifier trained on your calibration captures rather than on an average spine, plus a set of gates that hold back nudges when the geometry says the camera cannot read you reliably. We built those gates because we would rather miss a slouch than invent one.
Frequently Asked Questions
How accurate is webcam posture detection overall? For classifying seated posture with your upper body clearly visible, published systems land anywhere from about 70% to 98% depending on the posture and the camera angle, with forward lean at the top of that range. For measuring joint positions in absolute terms, error runs to 7 to 12 centimetres against motion capture. The first number is what a posture app needs. The second is what a physiotherapy lab needs.
Is a webcam as accurate as a wearable posture sensor? They fail differently. A wearable measures your torso angle directly and is unaffected by lighting or camera position, but it only knows about the one spot it is stuck to and needs charging and remembering. A webcam sees your whole upper body and needs nothing worn, but depends on a stable camera view. For comparing the two approaches properly, the deciding factor is usually whether you will actually keep using it.
Why does my posture app nag me when I am sitting fine? Almost always camera angle, calibration or lighting, in that order. An off-axis or steeply angled camera distorts the measurements the app relies on. Calibration done in an unnaturally good posture makes your normal sitting look like slouching. Dim lighting produces sensor noise, and noise makes keypoints jitter frame to frame. Fix those three and most false positives disappear.
Does the model get more accurate over time? Not on its own. The pose model is fixed and does not learn from you. What improves is the classifier built from your calibration, so recalibrating after a desk change, a new chair or a seasonal lighting shift is the closest thing to an accuracy upgrade you get.
Can a webcam tell if I am crossing my legs or tilting my pelvis? Effectively no. In the PoseTrack testing, crossed legs were detected in 0% of videos because the desk and arms block the knees. Anything below the desk edge should be treated as outside what a monitor-mounted camera can see.
Do more keypoints mean better accuracy? Not for desk posture. BlazePose returns 33 landmarks and MoveNet returns 17, but the ones that carry the signal for slouching are the ears, eyes, shoulders and hips. Extra hand and foot landmarks add nothing when you are sitting at a keyboard.
The Takeaway
Webcam posture detection sits a long way below a motion capture lab and a fair way above a colleague glancing over at you. The 2025 Scientific Reports work shows both halves of that in one paper. None of the models tested reached clinical accuracy. Most of the 3D ones still beat human raters doing visual assessment of knee flexion angles, and those raters were experienced.
Three things worth doing with all that.
Run the ten-minute test above this week on whatever app you already have. The false-alarm count over thirty quiet minutes of good sitting is the number that predicts whether you will still be using it in a month, and almost nobody measures it.
Before trusting any published figure, ask what it was measured against and on which posture. A bare percentage with no method behind it tells you nothing. That widely repeated “2-5 centimetres” claim is off by a factor of several against what the research reports.
And recalibrate whenever something changes. New desk, new chair, darker mornings, a bulkier jumper. For a calibrated system, accuracy is mostly a question of how honest the calibration was.
If you want the mechanism rather than the measurements, our guide to how AI posture detection works walks through the pipeline step by step. If you are choosing between apps, the posture app comparison covers what each one actually does with the camera. And if you would rather just try one, SitApp runs MoveNet locally on Mac, Windows and Linux, calibrates to your posture rather than to an average, and stays quiet when it cannot see you properly.