Lab · Geometry

Moving the camera or moving the object?

A point can acquire new coordinates without moving in the world. Follow four landmarks through a change of frame, then inspect what happens when the camera transform is used in the wrong direction.

What moves?

01 The world

xyzOCameraABCD

A, B, C and D are the same four vertices in every view. The outlined frustum marks the camera’s field of view; the axes stay fixed in the world.

Object0° yaw 0.00 m shift

Camera0° yaw 0.00 m shift

Try this: move the camera right. Which coordinates remain unchanged?

02 The camera image

360 × 240 px

Let C map camera coordinates to the world, and p be a world point. The image uses C−1p.

Correct · C⁻¹ p
uvABCD

All four landmarks in view.

03 The same landmarks

Coordinates (x, y, z), in metres. Camera values always use the correct inverse.
PointWorldCamera
A-0.75-0.65-2.50-0.75-0.65-3.50
B0.90-0.65-2.750.90-0.65-3.75
C-0.25-0.65-3.80-0.25-0.65-4.80
D0.100.95-3.100.100.95-4.10

The object has yaw 0° and a world-x shift of 0.00 m. Changing its pose changes both its world and camera coordinates; the camera stays fixed. Yaw turns the object about its own center.

The object has yaw 0° and a world-x shift of 0.00 m. Changing its pose changes both its world and camera coordinates; the camera stays fixed. Yaw turns the object about its own center.

Conventions and the calculation

World axes are right-handed, with +y up. The camera’s local +x points right, +y up, and −z forward. Positive yaw follows the right-hand rule about +y. The object turns about (0, 0, −3) m; the camera turns in place, initially at (0, 0, 1) m. Each mode retains its own pose.

The camera pose C is a 4 × 4 homogeneous matrix: its upper blocks are [R | t] and its last row is (0, 0, 0, 1). Here R is the 3 × 3 rotation and t is the camera’s world position. Matrix multiplication appends a final 1 to each point. In three coordinates, the inverse gives RT(p − t). For camera coordinates (x, y, z), depth d = −z, and image coordinates are u = cx + fx/d and v = cy − fy/d.

Focal length f = 220 px; principal point (cx, cy) = (180, 120) px. Image u runs right and v down. All vertices are labeled, including those a solid face would occlude. Points outside the image are reported; points on or behind the camera are not projected. No lens distortion, physics or measured imagery is modeled.

Let T be a rigid motion expressed in the world frame. Moving the camera from C to TCgives (TC)−1p = C−1T−1p. The same camera coordinates result from holding the camera fixed and moving the object by the full inverse motion T−1, including its translation.

Blender coordinates convert as (x, y, z) → (x, z, −y). Read the coordinates and conventions or the asset provenance.