Usage¶
Everything goes through one Detector (it loads its
models once and is safe to share between threads) and, usually, one call to
pose. The steps below show what that call does, and how
to do each step yourself.
1. Find the people¶
from PIL import Image
from poseorbit import Detector
detector = Detector() # or Detector(weights_dir=Path("weights"))
picture = Image.open("reference.png")
people = detector.detect(picture) # left to right
for index, person in enumerate(people):
print(index, person.bbox, person.visible)
Each Person has 133 COCO-WholeBody keypoints in the
picture's pixels, their scores, a bbox around what would be drawn, and
visible, the number of body joints (of 17) seen.
People come back left to right by the centre of the figure, so the same
picture numbers its people the same way every time: index 1 is the same
person on every call. A detection with fewer than 8 of 17 body joints seen is
not a person (a blank canvas otherwise "finds" one); with nobody left,
detect raises NoPersonError.
Pass depth=True to also get each keypoint's depth (Person.depth, in the
same pixels as x and y, positive away from the viewer). It costs about 0.3 s
more; pose asks for it only when the camera turns.
2. Choose whose pose¶
from poseorbit import ALL_PEOPLE, pose
pose(detector, picture, size=(832, 1216), person=0) # the leftmost
pose(detector, picture, size=(832, 1216), person=ALL_PEOPLE) # everyone (-1)
pose(detector, picture, size=(832, 1216)) # the most confident
Everyone together turns around one shared centre, so a group keeps its layout.
3. Turn it¶
from poseorbit import Camera
result = pose(detector, picture, size=(832, 1216), camera=Camera(yaw=-30, pitch=10))
Camera orbits the figure: yaw swings it to the
viewer's right (negative: left), pitch raises it. It is orthographic, so
turning never makes the figure bigger or smaller. Angles are held to ±90°
yaw and ±45° pitch: depth comes from a single picture, and past a side view
which limb is in front gets unreliable.
Front or back?
Some models do not tell front from back by the skeleton alone, and draw
a strongly turned figure from behind. Turn away from the side the figure
already shows (the nearer shoulder has the smaller Person.depth), keep
the turn moderate, and put "front view" in the prompt and "from behind"
in the negative.
4. Frame it¶
from poseorbit import Framing
# The face of a full-body picture: centre on it and zoom about five times.
result = pose(detector, picture, size=(832, 1216),
framing=Framing(zoom=5, x=0.62, y=0.2))
print(result.joints_in_frame) # body joints still inside the canvas
Framing crops the output canvas like a photo: the
canvas point (x, y, as fractions) moves to the middle and everything
scales by zoom (0.5 to 6). Below 1 it pulls back, leaving black around the
figure. Framing never moves the point the camera orbits.
joints_in_frame tells how much of the body is left: with a body-only
skeleton (the openpose style), a close-up keeps few points to follow.
5. Draw it¶
pose draws for you, in style="dwpose" or style="openpose"; see
Skeleton styles. The result's skeleton is a PIL image at
size, on black.
To draw keypoints of your own, call render; to fit a
skeleton drawn at another size onto your output, letterbox.
With diffusers¶
import torch
from diffusers import ControlNetModel, StableDiffusionXLControlNetPipeline
controlnet = ControlNetModel.from_pretrained(
"xinsir/controlnet-openpose-sdxl-1.0", torch_dtype=torch.float16)
pipe = StableDiffusionXLControlNetPipeline.from_pretrained(
"stabilityai/stable-diffusion-xl-base-1.0", controlnet=controlnet,
torch_dtype=torch.float16).to("cuda")
skeleton = pose(detector, picture, size=(832, 1216),
camera=Camera(yaw=45), style="openpose").skeleton
image = pipe("1girl, full body", image=skeleton, width=832, height=1216).images[0]
The skeleton must be drawn at the size you generate at, as here.
From another program¶
Run python -m poseorbit serve and post to /api/pose; see the
HTTP API. A generation server can embed the same
handler, poseorbit.api.handle, so both speak one
API.