What can a single camera actually tell us about motion?
I wanted to see how much real motion data I could get out of a plain golf swing video. No sensors on the body, no second camera, no depth camera. Just one video and a MediaPipe pose model that finds 33 body landmarks in every frame.
Landmarks are only the starting point. The end goal is a 3D model of the swing that a robotics simulator like MuJoCo can replay. That needs real physical state, and the catch is that a camera sees a flat image while a golf swing is mostly rotation.
Face-on camera. One video, turned into a synced pose scrubber. It shows the real camera angle plus an estimate of the down-the-line view. No mocap, no depth sensor.
The same swing filmed down-the-line. Compare each estimated view with the other clip’s real camera to see where the depth guess drifts. Source footage: Rory McIlroy’s Powerful Driver Swing by TaylorMade Golf.
- Left: the source video with MediaPipe landmarks drawn on top.
- Middle: the same pose as the camera saw it, in x and y.
- Right: the other angle, estimated from MediaPipe’s own depth signal.
Video In, State Out
The pipeline has five stages: pose detection → smoothing → phase detection → metrics and state → export.
frames = detect(video) # MediaPipe pose landmarker, every frame
frames = smooth(frames) # Savitzky-Golay over each landmark's trajectory
phases = detect_phases(frames, swing_type)
metrics = compute_metrics(frames)
states = compute_states(frames, metrics, phases)
export_states(states, output)
Each landmark comes with x, y, a rough depth z, and a visibility score. Raw, they jitter. A wrist shifts between frames even when the golfer is standing still, and every angle computed later inherits that noise. So before any math, a Savitzky-Golay filter runs over each landmark’s path through time (each frame is replaced by a parabola fitted through it and its 3 neighbors on each side). It removes the jitter without flattening the fast parts of the swing the way a plain moving average would.
What can we actually measure?
A camera gives us x and y. MediaPipe adds a z, but that’s the network’s guess at depth, not a measurement. So the real question is how much 3D rotation you can recover from a flat image.
Start with the hip line. When the hips turn, the line between the two hip landmarks looks shorter on screen. A rigid segment rotated by θ projects to its true length times cos(θ). We don’t know the true length, but at address the hips are roughly square to a face-on camera, so the address width works as a stand-in:
In code it’s three lines:
def rotation_from_span(span, address_span):
ratio = max(0.0, min(1.0, span / address_span))
return math.degrees(math.acos(ratio))
That’s the whole trick. It assumes the segment’s length and its distance from the camera stay constant through the swing. Golfers don’t fully cooperate, and we’ll get to where that breaks.
A second opinion from z
There’s another way to get the same angle. Instead of measuring how much the line shrank, look at it from above. MediaPipe’s z gives each landmark a depth, so the hip line has a direction in the x/z plane. Rotation is how far that direction has turned since address:
Same physical quantity, two different assumptions. The chart below plots both for hips and shoulders at every swing phase:
- Span trick: uses only x. Assumes constant segment length and camera distance.
- Z-derived: uses x and z. Assumes MediaPipe’s depth guess is reasonable.
All numbers in this post come from the face-on camera, and that’s on purpose. The span trick needs it. From down-the-line, the hip line points almost straight at the camera, so turning makes it look wider, not narrower. The trick only reads shrinkage, so it clamps to zero. On Rory’s down-the-line clip, it reads 0° for hips and shoulders all the way to the top.
Here’s what the two methods say on a face-on slow-motion clip of Rory McIlroy’s driver swing, in degrees from address:
| Phase | Hips (span) | Hips (z) | Shoulders (span) | Shoulders (z) |
|---|---|---|---|---|
| Takeaway | 0.0 | 1.1 | 34.3 | 24.3 |
| Backswing | 15.7 | 25.2 | 57.7 | 44.2 |
| Top | 25.9 | 40.3 | 52.5 | 43.9 |
| Downswing | 28.0 | 42.6 | 49.2 | 41.7 |
| Impact | 19.4 | 16.8 | 22.4 | 19.0 |
| Finish | 87.8 | 114.8 | 57.8 | 138.5 |
At impact the two agree within about 3° for both hips and shoulders. At the top and in the downswing they split, and not in the same direction. The z method sees about 15° more hip turn, while the span trick sees about 8° more shoulder turn. I don’t know yet which one is closer to the truth. The down-the-line clip helps as a visual check, but settling it needs two synced, calibrated cameras to triangulate the hip line in 3D, or mocap.
After impact they fall apart, and that part is expected. The span trick can’t go past 90°. Once the hips open past square, the line starts widening again, so it reads less turn while the body keeps rotating. The z method has no such ceiling.
Agreement isn’t ground truth. Both numbers come from the same pose model, so a shared error can fool both. But they fail in different ways. When they line up, I trust the number more. When they split, one of the assumptions just broke.
The view from above
The hip and shoulder lines at each swing phase, seen from above in the x/z plane. Color runs from dark purple at address to yellow at finish.
A line lying flat along x is square to the face-on camera. The more it tilts toward the depth axis, the more the body has turned. The angle between the address line and the top-of-backswing line is what the z method measures. How short each line gets along x is what the span trick reads. Comparing the two panels shows how much further the shoulders turn than the hips.
Where the trick breaks
The span trick reads every bit of on-screen shrinkage as rotation. So anything else that makes the hip or shoulder line look narrower gets counted as turn.
- Moving toward or away from the camera. Step back and everything looks smaller, the hip line included. The math reads that as extra rotation.
- Tilt. Span only measures width in x. The shoulders tilt a lot during a swing, and a tilted line is narrower in x even with zero turn. Some of that tilt gets counted as shoulder rotation.
- Jitter near address. This one comes from the math itself.
The acos curve is nearly vertical close to 1, and the ratio is 1 at address, exactly where the swing starts. In numbers, a 1% drop in span from address reads as about 8° of rotation. The same 1% drop at a ratio of 0.5 adds less than 1°. So small landmark noise early in the swing turns into large angle noise, and the span trick’s takeaway numbers are the ones to trust least. That’s my best guess for Rory’s takeaway row, where the span trick reads 34° of shoulder turn and z reads 24°.
From a single frame to a trajectory
So far everything is a measurement per frame. A simulator needs more than that. It needs state, and how that state changes over time. That’s what SwingState is for.
@dataclass
class SwingState:
frame_idx: int
timestamp: float
phase: str | None
pelvis_rotation: float | None # degrees from address (span trick)
torso_rotation: float | None
lead_arm_angle: float | None # degrees at the lead elbow
wrist_position: tuple[float, float, float] | None
head_position: tuple[float, float, float] | None
com_proxy: tuple[float, float, float] | None
angular_velocities: AngularVelocities
Each frame is one state, and the swing is the sequence of states. Phase and velocity are the two fields a single frame can’t give you. Both need the whole sequence. Velocity is just the change in angle between neighboring frames, divided by the time between them.
Differencing amplifies noise, so the velocity rows are the jumpiest part of the chart. The Savitzky-Golay pass is what keeps them readable.
This is what the sequence looks like traced out. It shows the paths of the lead wrist and a center-of-mass proxy over one swing, as the camera saw them. The COM proxy is a weighted blend of hips, shoulders and head, not a real center of mass. A real one needs segment masses, which a video doesn’t give you.
From script to command line
The whole pipeline sits behind two commands. export writes the state trajectory to JSON or CSV. report builds an interactive HTML page with every chart in this post, plus the scrubber from the clips at the top.
uv run python cli.py export --face-on swing.mp4 --club driver --output state_trajectory.json
uv run python cli.py report --face-on swing.mp4 --club driver --output report.html
Here are two frames from the export, at the top of Rory’s backswing, with values rounded.
[
{
"frame_idx": 241,
"timestamp": 10.052,
"phase": "top",
"pelvis_rotation": 25.878,
"torso_rotation": 52.538,
"lead_arm_angle": 127.266,
"wrist_position": [0.412, 0.39, -0.442],
"head_position": [0.444, 0.39, -0.168],
"com_proxy": [0.476, 0.496, -0.051],
"angular_velocities": {"pelvis": 21.607, "torso": 11.21, "lead_arm": 133.552}
},
{
"frame_idx": 242,
"timestamp": 10.093,
"phase": "top",
"pelvis_rotation": 26.564,
"torso_rotation": 52.631,
"lead_arm_angle": 128.601,
"wrist_position": [0.415, 0.389, -0.381],
"head_position": [0.441, 0.392, -0.162],
"com_proxy": [0.476, 0.497, -0.046],
"angular_velocities": {"pelvis": 16.444, "torso": 2.239, "lead_arm": 31.999}
}
]
The clip is slow motion, so timestamps and velocities are in video time, not real time. One state per frame, plain numbers, no golf logic baked in. That’s the format a simulator or a learning model can pick up directly.
Why this matters beyond golf
Golf is just the test case. Underneath is the chain any robot that learns from humans needs: sensor -> observation -> representation -> model -> action. This post covers the first three steps, from a camera to MediaPipe landmarks to a state trajectory with its weak spots spelled out. DeepMimic and SFV (Peng et al., 2018) showed that simulated characters can learn physical skills from motion clips, even from YouTube videos. The hard part is getting state you can trust out of a cheap camera. That’s the part I’m building.
So what can a single camera tell us about motion? More than I expected, as long as you know where the trick breaks. The code is on GitHub if you want to run it on your own swing.