Skip to content

Python API

Everything below is importable from poseorbit unless its heading says otherwise.

Detecting, choosing, turning and drawing

pose

pose(
    detector: Detector,
    image: Image,
    size: tuple[int, int],
    person: int | None = None,
    camera: Camera | None = None,
    style: Style = "dwpose",
    depth: bool = False,
    framing: Framing | None = None,
) -> PoseResult

Detect, choose, turn, frame and draw in one call.

A camera other than the front view needs depth, so it is detected then whatever depth says.

Parameters:

Name Type Description Default
detector Detector

A Detector; load it once and reuse it.

required
image Image

The reference picture.

required
size tuple[int, int]

The output's (width, height). The skeleton is drawn at it, the picture letterboxed onto it.

required
person int | None

An index into the left-to-right order, ALL_PEOPLE (-1) for everyone, or None for the most confident.

None
camera Camera | None

Where the camera stands; None for the front view.

None
style Style

"dwpose" or "openpose"; see render.

'dwpose'
depth bool

Detect depth even for the front view (to read Person.depth from the result).

False
framing Framing | None

Zoom and centre on the canvas; None for the whole picture.

None

Returns:

Type Description
PoseResult

The skeleton, everyone found and how it was drawn.

Raises:

Type Description
NoPersonError

Nobody in the picture is seen well enough.

ValueError

person is past the end of the people found.

PoseResult dataclass

PoseResult(
    skeleton: Image,
    people: list[Person],
    person: int,
    camera: Camera,
    framing: Framing,
    joints_in_frame: int,
)

What pose drew, and what it found on the way.

Attributes:

Name Type Description
skeleton Image

The skeleton on black, at the requested size.

people list[Person]

Everyone detected, left to right.

person int

Who was drawn: an index into people, or ALL_PEOPLE.

camera Camera

The camera actually used, after clamping.

framing Framing

The framing actually used, after clamping.

joints_in_frame int

Body joints (of 17) seen and inside the canvas, for the drawn person with the most. Few left (a close-up) means a body-only skeleton has little to follow.

Detector

Detector(
    weights_dir: Path | None = None, device: str = "cpu"
)

Finds people. Loads its models on first use; safe to share between threads.

Parameters:

Name Type Description Default
weights_dir Path | None

Where to keep the ONNX files (about 700 MB, downloaded on first use); default ~/.cache/poseorbit.

None
device str

onnxruntime's device, "cpu" by default.

'cpu'

detect

detect(image: Image, depth: bool = False) -> list[Person]

Everyone in image, ordered left to right by the centre of the figure, so the same picture always numbers its people the same way.

Parameters:

Name Type Description Default
image Image

The picture.

required
depth bool

Also estimate each keypoint's depth (RTMW3D-x, about 0.3 s more); see Person.depth.

False

Returns:

Type Description
list[Person]

The people found, left to right.

Raises:

Type Description
NoPersonError

Nobody is seen well enough (fewer than 8 of 17 body joints).

Person dataclass

Person(
    keypoints: ndarray,
    scores: ndarray,
    depth: ndarray | None = None,
)

One detected figure, in the picture's pixels.

visible property

visible: int

How many of the 17 body keypoints are seen: how sure this is a person.

bbox property

bbox: tuple[float, float, float, float]

(x0, y0, x1, y1) around the keypoints that would be drawn.

points_3d property

points_3d: ndarray

(133, 3): x, y and depth, all in pixels. Flat (depth 0) without depth.

NoPersonError

Bases: ValueError

The picture has nobody in it that is seen well enough to follow.

Camera and framing

Camera dataclass

Camera(yaw: float = 0.0, pitch: float = 0.0)

An orthographic camera orbiting the figure. Camera() is the front view.

Attributes:

Name Type Description
yaw float

Degrees; positive swings the camera to the viewer's right. Held to +-MAX_YAW (90).

pitch float

Degrees; positive raises the camera. Held to +-MAX_PITCH (45).

is_front property

is_front: bool

Whether this is the front view (no depth needed).

clamped

clamped() -> 'Camera'

The same camera held to MAX_YAW and MAX_PITCH.

Framing dataclass

Framing(zoom: float = 1.0, x: float = 0.5, y: float = 0.5)

Zoom and centre on the output canvas, like cropping a photo.

The canvas point (x, y) moves to the middle and everything scales by zoom around it. Framing() is the whole picture.

Attributes:

Name Type Description
zoom float

From MIN_ZOOM (0.5, room around the figure) to MAX_ZOOM (6, about a face close-up of a full-body picture).

x float

The canvas point to centre, as a fraction of its width.

y float

The same, as a fraction of its height.

is_whole property

is_whole: bool

Whether this is the whole picture: zoom 1, centred.

clamped

clamped() -> 'Framing'

The same framing held to the zoom limits and inside the canvas.

Drawing

render

render(
    keypoints: ndarray,
    scores: ndarray,
    size: tuple[int, int],
    style: Style = "dwpose",
) -> Image.Image

Skeletons on a black canvas, in the style a model was trained on.

Parameters:

Name Type Description Default
keypoints ndarray

(N, 133, 2) COCO-WholeBody keypoints in the canvas's pixels.

required
scores ndarray

(N, 133) their scores; points under KEYPOINT_THRESHOLD are not drawn.

required
size tuple[int, int]

The canvas's (width, height).

required
style Style

"dwpose" (rtmlib's COCO-WholeBody drawing: body, feet, hands and face) or "openpose" (the 18 body points, thick limbs that grow with the canvas, as xinsir's OpenPose ControlNet was trained).

'dwpose'

Returns:

Type Description
Image

The skeleton as an RGB picture.

Raises:

Type Description
ValueError

An unknown style.

letterbox

letterbox(
    skeleton: Image, size: tuple[int, int]
) -> Image.Image

A skeleton drawn for another size, scaled uniformly and centred on a black size canvas, so it lines up with the output however it was drawn.

Constants

Name Value Meaning
ALL_PEOPLE -1 person for everyone detected
MAX_YAW 90.0 The largest yaw, either way, in degrees
MAX_PITCH 45.0 The largest pitch, either way, in degrees
MIN_ZOOM / MAX_ZOOM 0.5 / 6.0 The framing's zoom range
KEYPOINT_THRESHOLD 0.3 A keypoint is seen (and drawn) at this score
STYLES ("dwpose", "openpose") The skeleton styles

Geometry

The steps pose takes, for drawing keypoints of your own.

view

view(
    points: ndarray, centre: ndarray, camera: Camera
) -> np.ndarray

Points seen from camera orbiting centre.

The front view (yaw 0, pitch 0) returns x and y unchanged.

Parameters:

Name Type Description Default
points ndarray

(..., 3) x right, y down, depth away from the viewer, in pixels.

required
centre ndarray

(3,) the point the camera orbits.

required
camera Camera

Where the camera stands.

required

Returns:

Type Description
ndarray

(..., 2) the points on screen, in the same pixels.

scene_centre

scene_centre(points: ndarray) -> np.ndarray

The point to orbit: the middle of everyone's hips, at the hips' depth.

points is (N, 133, 3). Depth is relative to each person's own hips, so every person turns in place; the centre's x and y are shared so a group keeps its arrangement.

fit

fit(
    keypoints: ndarray,
    source: tuple[int, int],
    size: tuple[int, int],
) -> np.ndarray

Keypoints from a source-sized picture placed on a size canvas, scaled uniformly and centred (letterboxed) -- never stretched, so a landscape reference does not come out as a squashed figure on a portrait canvas.

frame

frame(
    keypoints: ndarray,
    size: tuple[int, int],
    framing: Framing,
) -> np.ndarray

Canvas points after framing on a size canvas.

Parameters:

Name Type Description Default
keypoints ndarray

(..., 2) points on the canvas, in pixels.

required
size tuple[int, int]

The canvas's (width, height).

required
framing Framing

The zoom and centre.

required

Returns:

Type Description
ndarray

(..., 2) the framed points.

The HTTP handler

For a server of your own; see the HTTP API for the JSON.

handle

handle(
    detector: Detector,
    body: dict[str, Any],
    default_style: Style = "dwpose",
) -> dict[str, Any]

Answers one /api/pose request (the JSON above), without a web framework.

Parameters:

Name Type Description Default
detector Detector

The Detector to use.

required
body dict[str, Any]

The request's JSON object.

required
default_style Style

The style when the request names none; a generation server passes the one its loaded model needs.

'dwpose'

Returns:

Type Description
dict[str, Any]

The answer's JSON object.

Raises:

Type Description
BadRequest

The request is wrong; answer 400 with its message.

NoPersonError

Nobody in the picture; also a 400.

ValueError

A person index past the end; also a 400.

BadRequest

Bases: ValueError

The request is wrong; answer 400 with the message.

create_app

create_app(weights_dir: Path | None = None) -> 'FastAPI'

The FastAPI app answering /api/health and /api/pose.

Parameters:

Name Type Description Default
weights_dir Path | None

Passed to Detector.

None

Returns:

Type Description
'FastAPI'

A FastAPI application, for uvicorn or any ASGI server.