Skip to content

Releases: roboflow/supervision

supervision-0.30.8

Choose a tag to compare

@Borda Borda released this 06 Oct 19:12

v0.30.8 β€” Sharper video, cleaner labels

Video, YOLO labels, masks, VLM parsing and mAP all get more accurate.

  • VideoSink keeps OpenCV's video quality when OpenCV isn't installed.
  • from_yolo reads pose labels as boxes instead of polygons.
  • MeanAveragePrecision scores class-agnostic runs right when only one side has class IDs.
  • from_vlm returns one Florence-2 detection per object, not one per polygon.
  • from_inference masks no longer drift up to a pixel up and left.

Drop-in upgrade. Without OpenCV, videos get larger; YOLO labels with a negative width or height now raise ValueError.

✨ Spotlights / highlights

sv.VideoSink and sv.process_video keep quality without OpenCV

The PyAV fallback left the encoder bit rate unset, so mp4v and MJPG files came out at under half of what cv2.VideoWriter writes. It now uses OpenCV's rate settings for every codec except H.264. Files get larger and vp09 may encode more slowly; codec="avc1" keeps files small where an H.264 encoder is available. (#2661)

import supervision as sv

video_info = sv.VideoInfo.from_video_path("in.mp4")
with sv.VideoSink("out.mp4", video_info) as sink:  # OpenCV not installed
    for frame in sv.get_video_frames_generator("in.mp4"):
        sink.write_frame(frame)
# before: mp4v written at under half OpenCV's bit rate, visibly softer
# now:    same bit rate OpenCV's writer uses

YOLO pose labels load as boxes

A pose row is a box followed by keypoints. from_yolo used to parse the whole row as a polygon, giving wrong boxes and masks nobody asked for. It now reads the box and skips the keypoints that kpt_shape declares. (#2655)

ds = sv.DetectionDataset.from_yolo(
    images_directory_path="pose/images",
    annotations_directory_path="pose/labels",
    data_yaml_path="pose/data.yaml",  # kpt_shape: [17, 3]
)
# before: polygon-parsed boxes and masks
# now:    one box per row

Class-agnostic mAP with one-sided class IDs

A perfect match scored zero when only one side carried class IDs, such as SAM proposals checked against labeled ground truth. With class_agnostic=True, both sides now count as one class.

One Florence-2 detection per object (#2648)

Florence-2 returns a segmented object as a list of polygons, one per connected region. An object split in two used to come back as two detections; the polygons now merge into one mask with one box around all of them.

Roboflow masks sit on the right pixels (#2649)

from_inference truncated sub-pixel polygon vertices, shifting each mask up and left by up to a pixel. Vertices are now rounded, the way the COCO, YOLO, LabelMe and Pascal VOC loaders already do.

πŸ”„ Migration guide

No migration required for this release.

πŸ“ Notable changes

🌱 Changed

  • sv.DetectionDataset.from_yolo raises ValueError naming the annotation file when a label has a negative width or height; it used to load a box with x_min past x_max, which made Detections.area negative and skewed IoU and NMS. as_yolo now orders the corners of a reversed box before measuring, so it no longer writes a file the loader refuses. (#2663)

πŸ”§ Fixed

  • sv.VideoSink and sv.process_video write mp4v, MJPG and other non-H.264 video at OpenCV's bit rate when OpenCV isn't installed. A frame rate of zero or less now raises RuntimeError in sv.VideoSink, as it does with OpenCV. (#2661)
  • sv.DetectionDataset.from_yolo reads the box of Ultralytics pose labels and skips their keypoints; a kpt_shape other than [K, 2] or [K, 3] raises ValueError. (#2655)
  • sv.DetectionDataset.from_yolo names a malformed annotation line, one with too few values or, for OBB, not nine, in a ValueError instead of failing on an array shape. (#2665)
  • sv.metrics.MeanAveragePrecision(class_agnostic=True) treats detections without class IDs as the same class as labeled ones, and unsigned class ID arrays no longer raise OverflowError on NumPy 2. (#2650)
  • sv.Detections.from_vlm with sv.VLM.FLORENCE_2 merges the polygons of one instance into one detection for <REFERRING_EXPRESSION_SEGMENTATION> and <REGION_TO_SEGMENTATION>, and skips instances with no usable polygon. (#2648)
  • sv.Detections.from_vlm with sv.VLM.QWEN_2_5_VL or sv.VLM.QWEN_3_VL recovers complete detections from a response cut off inside a bbox_2d array or right after a complete object. (#2666)
  • sv.Detections.from_inference rounds polygon vertices to the nearest pixel before rasterising masks, and raises ValueError for NaN or infinite vertices. (#2649)
  • sv.xyxy_to_mask returns an empty mask for a box entirely left of or above the image when its maximum coordinate is a negative fraction. (#2646)
  • sv.Detections.get_anchors_coordinates computes axis-aligned midpoint anchors without integer overflow. (#2660)
  • sv.LineZone.trigger ages crossing history on frames whose detections lack tracker_id, so a reused track ID no longer creates a false crossing after the track expired. (#2644)
  • sv.LineZoneAnnotator(text_orient_to_line=True) no longer raises TypeError without OpenCV for lines drawn right to left. (#2659)
  • The Ultralytics, Inference and YOLO-NAS speed estimation examples measure elapsed time from frame indices; a vehicle missed in one frame of three was reported about 44% too fast. (#2654)
  • examples/speed_estimation/rfdetr_example.py no longer raises AttributeError: supervision._cv2 now provides getPerspectiveTransform and perspectiveTransform. (#2652)

πŸ† Contributors

  • Mohammad Hijjawi (@MohammadHijjawi97, LinkedIn) β€” fixed video quality without OpenCV, Florence-2 instance merging and Roboflow mask rounding.
  • Kari Pikkarainen (@kari-pikkarainen, LinkedIn) β€” fixed speed estimation timing, the NumPy flip fallback and the perspective-transform fallbacks.
  • Miral Amin (@aminmiral) β€” made YOLO loading reject negative extents and name malformed lines.
  • NIKHIL (@Nikhi00718) β€” fixed LineZone history expiry and anchor overflow.
  • kevin (@kevin9327) β€” made Qwen parsing recover from cut-off responses.
  • A Aswanth Raj (@aswanth-07, LinkedIn) β€” fixed class-agnostic mAP.
  • Devulapalli Naga Sri Vaishnavi (@Vaishnavi220506) β€” fixed masks for off-frame fractional boxes.
  • JANG BYUNGKUN (@8rulerstar) β€” fixed YOLO pose label loading.

Automated contributions: @dependabot


Full changelog: 0.30.7...0.30.8

supervision-0.30.7

Choose a tag to compare

@Borda Borda released this 04 Oct 12:41

v0.30.7 β€” Fewer hangs, truer mAP

Eight bug fixes: video processing, dataset export, mAP and Transformers loading no longer fail or hang silently.

  • process_video raises on a failed write instead of hanging.
  • MeanAveragePrecision matches pycocotools where recall lands on a threshold.
  • Dataset export can write back into the folder the images came from.
  • from_transformers reads semantic output that includes per-pixel scores.

Drop-in upgrade, no code changes. Stored mAP baselines may shift slightly.

✨ Spotlights / highlights

sv.process_video stops hanging (#2636)

A frame the writer rejects, or a callback that returns None, now raises instead of blocking forever.

import supervision as sv

def callback(frame, index):
    sv.BoxAnnotator().annotate(frame.copy(), detections)  # forgot `return`

sv.process_video("in.mp4", "out.mp4", callback)
# before: hangs after `writer_buffer` frames
# now:    TypeError naming the frame

sv.metrics.MeanAveragePrecision matches pycocotools (#2638)

Recall thresholds and recall are float64, as in COCOeval. A class whose recall lands exactly on a threshold used to score slightly high: mAP@50 0.7470 instead of 0.7415 in one case.

Datasets re-export in place (#2637)

ds = sv.DetectionDataset.from_coco(
    images_directory_path="data",
    annotations_path="data/_annotations.coco.json",
)
ds.as_coco(
    images_directory_path="data",
    annotations_path="data/_annotations.coco.json",
)
# before: shutil.SameFileError
# now:    annotations rewritten, images left in place

Transformers semantic output with scores (#2643)

sv.Detections.from_transformers accepts results from return_segmentation_scores=True instead of raising KeyError: 'segments_info'.

πŸ”„ Migration guide

No migration required for this release.

πŸ“ Notable changes

πŸ”§ Fixed

  • sv.process_video raises RuntimeError("Writer thread raised: ...") or TypeError instead of hanging when a frame cannot be written or a callback returns None. (#2636)
  • sv.metrics.MeanAveragePrecision computes IoU and recall thresholds, IoUs and recall in float64, matching pycocotools. mAP changes only for classes whose recall lands exactly on one of the 101 thresholds, by up to about 0.007 mAP@50. (#2638)
  • sv.DetectionDataset.as_yolo, as_pascal_voc, as_coco, as_createml and as_labelme export into the folder the images were loaded from instead of failing with shutil.SameFileError. (#2637)
  • sv.Detections.from_transformers accepts semantic segmentation output with segmentation_scores; per-pixel scores stay out of per-detection confidence. (#2643)
  • sv.crop_image clips finite crop coordinates outside the 32-bit integer range to the image bounds, instead of wrapping to an empty crop. (#2642)
  • sv.DetectionDataset.as_labelme gives disconnected components of one mask a shared group ID, so from_labelme rebuilds one detection. (#2640)
  • sv.DetectionDataset.as_labelme, as_yolo and as_pascal_voc export in-memory grayscale (height, width) images. (#2641)
  • sv.KeyPoints.with_nms keeps skeletons with zero joints instead of raising a zero-size reduction error. (#2639)

πŸ† Contributors

  • kevin (@kevin9327) β€” stopped process_video hangs, made mAP match pycocotools, and fixed in-place dataset export.
  • Marvel Harisson (@INo-xious, LinkedIn) β€” fixed LabelMe grouping, grayscale dataset export and keypoint NMS.
  • NIKHIL (@Nikhi00718) β€” fixed crop_image clipping and Transformers semantic output loading.

Full changelog: 0.30.6...0.30.7

supervision-0.30.6

Choose a tag to compare

@Borda Borda released this 29 Sep 11:07

0.30.6: Pillow-conversion, dataset-export, and line-crossing correctness fixes

supervision 0.30.6 is a patch release closing 12 library correctness bugs and landing 2 behavior refinements across datasets, annotators, key points, line-crossing counting, and the Pillow-based fallback backend, plus a new docs guide and a notebook dead-link fix. No new public API, no breaking changes. The broadest-reach fix rewrites pillow_to_cv2, the entry point every annotator except RichLabelAnnotator and several image helpers use to convert a Pillow image: it used to pass raw mode bytes straight through, so a 1-bit mask, a 16-bit depth map, an LA/PA image, or a CMYK JPEG all drew wrong or crashed. Two fixes close silent-wrong-data bugs rather than crashes: LineZone.trigger let unconfirmed tracks inflate crossing counts, and DetectionDataset.as_yolo/.as_pascal_voc silently dropped box-only objects from a mixed polygon/box COCO dataset. DetectionsSmoother no longer crashes when two tracked objects first seen on different frames carry different metadata (e.g. RF-DETR's source_image); its own docstring pipeline was broken. Every library fix ships a regression test.

✨ Spotlights / highlights

sv.pillow_to_cv2 now converts every Pillow mode the way cv2.imread would

Raw mode bytes used to pass straight through: a 1-bit mask came back as 0/1 instead of 0/255, a 16-bit depth map wrapped modulo 256 and drew as noise, an LA/PA image crashed in cvtColor on its 2-channel array, and a CMYK JPEG drew its cyan/magenta/yellow ink as red/green/blue. This function is the entry point for every annotator except RichLabelAnnotator (which stays on the Pillow path) plus crop_image, resize_image, letterbox_image, scale_image, tint_image, grayscale_image, and plot_image whenever handed a Pillow image, so the bug reached the whole drawing surface, not just one call site. A single-channel grayscale image also drew from a read-only buffer view before this fix. Annotating into it silently failed; it now hands back a writable copy.

from PIL import Image

depth_map = Image.open("depth.tif")  # mode "I;16"
scene = sv.pillow_to_cv2(depth_map)
# clipped to the 16-bit range and keeps its high byte, instead of wrapping mod 256

RGB, RGBA, grayscale, and palette images are unchanged.

sv.LineZone.trigger no longer lets unconfirmed tracks inflate counts

Detections with a negative tracker_id (how trackers like ByteTrackTracker mark an unconfirmed track) all keyed under one shared id. Two distinct unconfirmed objects crossing on opposite sides in the same frame read as one track oscillating, silently inflating in_count/out_count with crossings no confirmed track ever made. Confirmed tracks (tracker_id >= 0) are counted exactly as before.

sv.DetectionDataset.as_yolo / .as_pascal_voc stop dropping box-only objects

from_coco gives an all-zero mask to any annotation without its own segmentation when a sibling annotation has one, and to every annotation when force_masks=True. The exporters only ever wrote the polygon traced from a mask, so a mixed polygon/box COCO dataset silently lost its box-only objects on export, and a box-only COCO dataset loaded with force_masks=True wrote empty YOLO label files or object-less Pascal VOC files. A detection with an empty or contour-less mask now exports as its bounding box instead.

sv.DetectionsSmoother no longer crashes on a second tracked object

Each smoothed track was built on the oldest frame in its window, so two objects first seen on different frames carried different frames' metadata (exactly what the RF-DETR / inference connectors attach), and Detections.merge rejected the mismatch. class_id and data (e.g. class_name) now follow the current frame instead of lagging up to length - 1 frames behind a class change; xyxy, confidence, and oriented-box corners are still averaged as before.

Dataset splits reject an out-of-range ratio instead of silently mis-splitting

sv.DetectionDataset.split, sv.ClassificationDataset.split (both take split_ratio), and the internal train_test_split helper they call (takes train_ratio) never validated the ratio. A finite out-of-range value looked like a successful split: for 10 images, a ratio of -0.2 returned 8/2 via negative slicing, and 1.2 (or an accidental percentage like 80) returned 10/0 with no held-out data, no error either way.

train, test = dataset.split(split_ratio=80)  # meant 80%
# now raises ValueError naming the problem, instead of silently returning
# every image as train and none as test

0 and 1 keep their existing meanings. Not an API break: no signature changed, only previously-silent input now raises.

πŸ”„ Migration guide

No breaking changes in this release.

No deprecations or removals landed in 0.30.6 either. All scheduled remove_in markers in the codebase target 0.31.0 or 0.32.0 and are untouched by this patch release.

Two entries change output values rather than only fixing a crash or a silently-wrong count; not API breaks, but worth checking if your code depends on the old values:

  • sv.KeyPoints.from_ultralytics / .from_inference / .from_detectron2: as_detections().confidence now carries the model's own detection score, not the mean of the per-keypoint confidences.
  • sv.DetectionsSmoother: class_id and data fields on a smoothed track now follow the current frame instead of lagging behind a class change.

πŸ“ Notable changes

πŸ”§ Fixed

Annotators / key points

  • sv.VertexEllipseHaloAnnotator now draws the whole halo of a key point whose covariance ellipse is not horizontal. A vertical or diagonal ellipse used to be clipped to a thin band because the fade box was sized as if the major axis always ran along the image x axis. Horizontal ellipses are drawn exactly as before. (#2634)
  • sv.KeyPoints.from_ultralytics, .from_inference and .from_detectron2 now keep each object's detection score as detection_confidence instead of dropping it, which previously made sv.KeyPoints.with_nms raise ValueError and as_detections() report the mean keypoint confidence instead of the model's own score. (#2633)
  • sv.LabelAnnotator now sizes the label background correctly for a label containing a blank line. Text drew each blank line at full height while the background measured it as zero, pushing the last line outside the box. Labels without blank lines are unchanged. (#2632)
  • sv.IconAnnotator now draws a palette icon that carries its own alpha channel (Pillow PA mode, storable in TIFF) instead of raising ValueError. Palettes without alpha, and every read with OpenCV installed, are unchanged. (#2624)

Datasets

  • sv.DetectionDataset.as_yolo / .as_pascal_voc now write a detection whose mask is empty or has no valid contour as its bounding box instead of silently leaving it out of the label file. (#2631)
  • sv.DetectionDataset.from_yolo now loads a label row carrying a trailing confidence or tracker id (written by Ultralytics save_txt with save_conf=True or tracking on) instead of raising ValueError, for both box and segmentation rows. The extra token is ignored; box and polygon geometry are unchanged. An odd-length polygon row with no extra field now loses its last value rather than raising, since coordinate parity is the only signal separating the two cases. (#2619, #2626, #2635)

Tracking

  • sv.DetectionsSmoother no longer raises ValueError: Conflicting metadata when a second tracked object, first seen on a different frame, enters the smoothing window. (#2628)
  • sv.LineZone.trigger now ignores detections with a negative tracker_id (the value trackers such as ByteTrackTracker report for an unconfirmed track) instead of letting them silently inflate in_count/out_count. (#2623)

Detection utils

  • sv.polygon_to_mask now accepts a list/tuple/array-like of [x, y] vertices (not just a NumPy array) and returns an all-zero mask for an empty or under-3-vertex polygon instead of crashing inside OpenCV/NumPy with an opaque error. A malformed polygon now raises ValueError naming the problem. (#2622)
  • sv.Detections.from_sam3 now keeps SAM 3 PVS contour fragments with fewer than 3 vertices instead of dropping them. Single points and 2-point edges are rasterized directly into the mask and bounding box. (#2625)

Image / IO

  • sv.pillow_to_cv2 now converts every Pillow mode to the 8-bit array cv2.imread would produce, reaching every annotator plus crop_image, resize_image, letterbox_image, scale_image, tint_image, grayscale_image, and plot_image. (#2614)
  • sv.CSVSink now writes UTF-8 on every platform, preserving non-English detection labels and custom fields on Windows. (#2615)

Docs

  • Dead documentation links fixed across the published notebooks (quickstart.ipynb, annotate-video-with-detections.ipynb, underestand-visitors-with-yolo-world.ipynb). (#2630)

🌱 Changed

  • sv.DetectionDataset.split, `sv.ClassificationDatase...
Read more

supervision-0.30.5

Choose a tag to compare

@Borda Borda released this 22 Sep 12:22

0.30.5: Tracking, metrics, and image-drawing correctness fixes

supervision 0.30.5 is a patch release closing 11 correctness and crash bugs across line-crossing tracking, mAR@K scoring, model connectors, key points, and image drawing/loading β€” no new public API of note, no breaking changes. The most consequential fixes are silent, not crashes: LineZone.trigger miscounted crossings by one per flicker whenever a tracker briefly touched the far side of the line, and MeanAverageRecall scored mAR@K against the wrong predictions when a lower-ranked one fit a target more tightly than one within the top K. The remaining fixes close a hard crash in InferenceSlicer on conflicting slice metadata (plus two related Detections.__eq__ bugs), bring the OpenCV-free fallback backend to parity with OpenCV for rotated videos, CMYK images, transparent/1-bit PNGs, and default JPEG/WebP write quality, fix two drawing bugs that reproduce with OpenCV installed (16-bit draw_image, grayscale IconAnnotator icons), and fix two isolated bugs in KeyPoints.as_detections and plot_images_grid. Every fix ships a regression test.

✨ Spotlights / highlights

sv.LineZone.trigger no longer counts flicker as a crossing

A crossing was confirmed whenever the oldest entry of a minimum_crossing_threshold + 1 frame history differed from every later entry, which never verified the tracker had actually settled on the side it supposedly came from. With minimum_crossing_threshold=2 the side sequence A,A,A,B,A,A,A counted a crossing into A β€” the side the object never left β€” so counts drifted by one per flicker, in the wrong direction; three separated flickers gave in_count=3 instead of 0.

line_zone = sv.LineZone(
    start=sv.Point(0, 0), end=sv.Point(0, 100), minimum_crossing_threshold=2
)
# a tracker that flickers to the far side for one frame and back
# no longer registers a phantom crossing; only a sustained crossing counts

Crossings are now measured against the last side a tracker was confirmed on, not against the oldest history entry. Sustained crossings and minimum_crossing_threshold=1 (the default) are unchanged.

sv.metrics.MeanAverageRecall now scores mAR@K from each image's own top K

The matcher pairs predictions to targets by highest IoU, not confidence, so a prediction ranked below K could take a target away from one ranked within it. Adding a low-confidence duplicate that fit a target more tightly than the top prediction actually lowered mAR@1 β€” one target with a top prediction at IoU 0.71 scored mAR@1 0.5 alone but 0.0 once a second, unrelated prediction at IoU 1.0 and confidence 0.1 was added. Each detection limit now matches only its own top K predictions.

sv.InferenceSlicer no longer crashes when slices disagree on metadata

RF-DETR and inference-package connectors attach a source_image array per slice; any mismatch across slices previously crashed the merge outright. Any metadata key that isn't identical across every slice is now dropped with a one-time warning naming the dropped keys; source_image is reattached afterward as the full input image.

slicer = sv.InferenceSlicer(callback=callback)
detections = slicer(image)  # metadata.source_image is the full input image again

Fixing this also closed two Detections.__eq__ bugs: NaN in a float-array metadata value now compares equal to itself, and a list-valued value against an ndarray-valued one now returns False instead of raising ValueError.

Rotated videos read upright without OpenCV

OpenCV's FFmpeg backend applies a video's display-rotation matrix to decoded frames and to the reported width/height; the OpenCV-free fallback ignored it, so a portrait phone video came back sideways with its width/height swapped.

sv.KeyPoints.as_detections drops invisible key points from the derived box

as_detections stretched each box over every key point that wasn't [0, 0] or non-finite, ignoring visible β€” unlike with_nms, which already respected it. A pose whose low-confidence joints were predicted off-frame no longer produces an inflated box.

πŸ”„ Migration guide

No breaking changes in this release.

No deprecations or removals landed in 0.30.5 either. All scheduled remove_in markers in the codebase target 0.31.0 and are untouched by this patch release.

sv.config.SOURCE_IMAGE_METADATA_FIELD is a new public constant (added alongside the InferenceSlicer fix) naming the Detections.metadata key RF-DETR/inference-package connectors use for the source image β€” additive, no existing code needs to change.

πŸ“ Notable changes

πŸ”§ Fixed

Tracking / metrics

  • sv.LineZone.trigger no longer counts a spurious crossing in the opposite direction when a tracker flickers to the far side of the line for fewer than minimum_crossing_threshold frames. Crossings are now measured against the last side a tracker was confirmed on, and once a tracker has a confirmed side, a new one replaces it only after being held for the full threshold. Sustained crossings, minimum_crossing_threshold=1 (the default), and per-tracker isolation are unchanged. (#2600)
  • sv.metrics.MeanAverageRecall now scores mAR@K from each image's K most confident predictions alone, instead of matching every prediction first and only then keeping the top K by confidence. mAR@1 and mAR@10 now equal the recall of the top 1 and top 10 predictions per image, as documented; mAR@100 changes only for images with more than 100 predictions. (#2604)

Model connectors

  • sv.InferenceSlicer no longer raises when slices disagree on metadata (e.g. a source_image NumPy array attached per-slice by RF-DETR/inference-package connectors), which previously crashed the merge outright. Any metadata key that isn't identical across every slice is now dropped from the merged result, with a SupervisionWarnings warning naming the dropped keys, emitted once per slicer instance. source_image is a special case: removed from each slice before the lenient merge, then reattached afterward as the full input image. Fixing this also closed two related Detections.__eq__ bugs: metadata holding NaN in a float array now compares equal to itself instead of always reading unequal, and comparing a list-valued metadata value against an ndarray-valued one now returns False instead of raising ValueError. (#2596)

OpenCV-free fallback backend

  • sv.VideoInfo.from_video_path, sv.get_video_frames_generator and sv.process_video now turn a rotated video upright when OpenCV is not installed, as they already do with OpenCV. The fallback now applies quarter and half turns, the same angles OpenCV applies, using the container's display-rotation matrix. Videos without a display rotation, and every read with OpenCV installed, are unchanged. (#2601)
  • sv.IconAnnotator and sv.draw_image now draw CMYK JPEG and TIFF images in their colors when OpenCV is not installed, instead of returning the four ink channels as if they were blue/green/red/alpha. Other images, and every read with OpenCV installed, are unchanged. (#2602)
  • sv.ImageSink and the dataset exports that encode in-memory images now write JPEG and WebP files at OpenCV's default quality (JPEG 95, lossless WebP) when OpenCV is not installed, instead of Pillow's defaults (JPEG 75, lossy WebP). The fallback's in-memory encoder also now accepts .jpe, .tif, .jp2, and .pgm, which it previously rejected even though file writes already accepted them. PNG and every write with OpenCV installed are unchanged. (#2592)
  • sv.IconAnnotator and sv.draw_image now keep the transparency of grayscale PNGs with alpha, RGB PNGs with a transparent color, and 1-bit PNGs, when OpenCV is not installed β€” the fallback previously returned Pillow's pixel layout for IMREAD_UNCHANGED rather than OpenCV's, causing shape-mismatch ValueErrors or black-drawn transparent pixels. Other images, and every read with OpenCV installed, are unchanged. (#2588)

Annotators / drawing

  • sv.draw_image now scales a 16-bit PNG down to 8 bits on load, instead of blending it into an 8-bit scene and clipping channel values above 255 β€” this reproduces with OpenCV installed, since IMREAD_UNCHANGED keeps a 16-bit PNG at 16 bits. Eight-bit images, and images passed as arrays, are drawn as before. (#2603)
  • sv.IconAnnotator now draws grayscale PNG icons instead of failing with IndexError β€” the overlay only handled BGR and BGRA arrays, and cv2.imread(..., IMREAD_UNCHANGED) returns a 2-D array for a grayscale PNG without alpha, with or without OpenCV installed. A grayscale icon is now expanded to BGR on load, and a 16-bit icon is scaled to 8 bits, as cv2.imread does by default. Eight-bit color icons are drawn as before. (#2591)

Key points / utils

  • sv.KeyPoints.as_detections now leaves key points marked not visible out of the box it derives for each skeleton, as sv.KeyPoints.with_nms already does. A skeleton with no visible key point is now dropped, as one with only missing key points already was. Key points without visible convert as before. (#2605)
  • sv.plot_images_grid now plots a single image in a grid_size=(1, 1) grid, instead of failing with AttributeError β€” plt.subplots returns a lone Axes rather than an array for a 1x1 grid. Grids with more than one cell plot...
Read more

supervision-0.30.4

Choose a tag to compare

@Borda Borda released this 17 Sep 15:12

0.30.4: Dataset, connector, and key point correctness fixes

supervision 0.30.4 is a patch release closing 14 correctness and crash bugs across dataset loaders/exporters, model connectors, and key point/annotator/video handling β€” no new public API, no breaking changes. The most consequential fixes are silent, not crashes: COCO polygon masks loaded shifted by up to a pixel, EXIF-rotated photos loaded with swapped width/height across two loaders and both image backends, and non-ASCII class names were mangled or rejected on Windows across four loaders and one writer. The remaining fixes close hard crashes in the Transformers v4/v5 connectors, Ultralytics pose loading, TraceAnnotator, iterative_seek, and three more dataset-format edge cases. Every fix ships a regression test.

✨ Spotlights / highlights

sv.DetectionDataset.from_coco no longer shifts mask polygons by up to a pixel

COCO polygons commonly hold sub-pixel float coordinates, but the loader cast them straight to int32, which truncates rather than rounds. Every polygon mask loaded shifted up and to the left by up to a pixel β€” the same polygon loaded one pixel apart depending on the format it was stored in (a square with corners at 2.6/7.6 covered pixels 2–7 from COCO but 3–8 from LabelMe, an IoU of 0.53 between the two masks). Silent, not a crash β€” training data was quietly misaligned.

dataset = sv.DetectionDataset.from_coco(
    images_directory_path="images",
    annotations_path="annotations.json",
)  # polygon vertices now rounded to nearest pixel, matching from_yolo/from_labelme/from_pascal_voc

A non-finite vertex now raises ValueError naming the annotation id, instead of silently propagating.

from_yolo/as_coco no longer swap width and height on EXIF-rotated photos

Photos from phones are often stored sideways with an EXIF orientation tag. cv2.imread applies that tag, but the size read used Pillow's file-header read, which doesn't. For a quarter-turned photo, from_yolo scaled normalized boxes and polygons by the swapped width/height, so they landed outside the image, and as_coco wrote the swapped dimensions. The OpenCV-free fallback backend now also applies the orientation tag, matching how OpenCV itself handles every read except IMREAD_UNCHANGED β€” the same file previously loaded with a different shape depending on whether opencv-python was installed.

Dataset loaders/writers no longer mangle non-ASCII class names on Windows

from_coco, from_labelme, from_createml, from_yolo read JSON/YAML, and as_pascal_voc wrote XML, using the platform's default encoding β€” cp1252 on Windows. A class name like cafΓ© loaded as café; 고양이 raised UnicodeDecodeError. All five now read/write UTF-8 explicitly.

sv.Detections.from_transformers now loads Mask2Former/MaskFormer overlap-safe binary maps

post_process_instance_segmentation(return_binary_maps=True) β€” the option Transformers recommends when instances can overlap β€” returns a (num_instances, H, W) stack of binary maps, but the v5 instance path compared it against each segment's id as if it were an id-map, producing a 4-D array mask_to_xyxy rejected outright. Each segment now indexes the stack at its own id, so overlapping instances keep their full masks.

detections = sv.Detections.from_transformers(
    transformers_results=processor.post_process_instance_segmentation(
        result, target_sizes=[image.size[::-1]], return_binary_maps=True
    )[0],
    id2label=model.config.id2label,
)  # overlapping-instance results now load instead of crashing

sv.DetectionDataset.from_pascal_voc no longer silently drops bmp/tif/webp images

The loader only listed .jpg, .jpeg, .png, so every other image was left out without a warning β€” even though from_yolo/from_folder_structure load those formats and as_pascal_voc writes annotations for them. A dataset exported to Pascal VOC and read back came back smaller than it went out.

πŸ”„ Migration guide

No breaking changes in this release.

No deprecations or removals landed in 0.30.4 either. Eight scheduled-removal commits (ByteTrack, supervision.keypoint, create_tiles/overlay_image, keypoint validators, the legacy MeanAveragePrecision, LMM/from_lmm, and others) plus a new MetricResult ABC exist on develop but are not part of this cherry-picked patch release β€” they target 0.31.0 and will get their own migration guide when that release ships.

Two internal (non-underscored but unexported) helper functions changed signature as part of fixes in this release: parse_polygon_points (dataset/formats/pascal_voc.py) now returns float64 instead of int, and detections_from_xml_obj's xyxy is now list[list[float]] instead of list[list[int]]. Neither is re-exported from supervision.__init__ or supervision.dataset.__init__, so this is not a public API change β€” no action needed unless you import these paths directly.

πŸ“ Notable changes

πŸ”§ Fixed

Dataset loaders / exporters

  • sv.DetectionDataset.from_coco now rounds polygon vertices to the nearest pixel before rasterising masks, instead of truncating them to int32, as from_yolo/from_labelme/from_pascal_voc already do. Truncation shifted every mask up and to the left by up to a pixel β€” a polygon with corners at 2.6/7.6 covered pixels 2 to 7 from COCO but 3 to 8 from LabelMe, an IoU of 0.53 between the two masks. A non-finite vertex now raises ValueError naming the annotation id; integer vertices and RLE masks load as before. (#2587)
  • sv.DetectionDataset.from_coco/from_labelme/from_createml/from_yolo now read JSON/YAML as UTF-8, and as_pascal_voc writes XML as UTF-8, instead of the platform default (cp1252 on Windows). A COCO category or YOLO data.yaml name outside ASCII broke on Windows: cafΓ© loaded as café, 고양이 failed with UnicodeDecodeError, and as_pascal_voc either failed with UnicodeEncodeError or wrote a file from_pascal_voc rejected with ParseError: not well-formed (invalid token) β€” Linux/macOS, already UTF-8 by default, are unaffected. (#2585)
  • sv.ClassificationDataset.as_folder_structure now copies source image files unchanged instead of round-tripping them through cv2.imread/cv2.imwrite, which dropped alpha channels, downcast 16-bit PNGs to 8-bit, and recompressed JPEGs on every export β€” matching sv.DetectionDataset's exports (as_yolo/as_pascal_voc/as_coco), which already copied; exporting into the folder the dataset was loaded from previously rewrote its own images the same lossy way. Images held in memory are still encoded with cv2.imwrite. (#2584)
  • sv.DetectionDataset.from_labelme now finds images for LabelMe files saved on Windows when loading on Linux or macOS. imagePath is written with \, but the loader took only Path(...).name, which doesn't split on \ on POSIX β€” the whole value became the filename and reading failed with ValueError: Could not read image from path. The filename is now taken with either separator on every system, as LabelMe itself does when reading its own files; forward-slash paths, and every file on Windows, are unchanged. (#2581)
  • sv.DetectionDataset.from_yolo now loads label files whose class ids are written as decimals (e.g. 1.0 0.5 0.5 0.2 0.4), which previously aborted the whole load with ValueError: invalid literal for int() β€” np.savetxt writes floats by default and Ultralytics tolerates them. Whole numbers load in any notation now; fractional, non-finite, or non-numeric ids still raise, naming the offending id. (#2580)
  • sv.DetectionDataset.from_yolo/as_coco now size EXIF-oriented images the way cv2.imread loads them (swapping width/height for orientations 5 to 8), instead of reading the un-rotated file-header size via Pillow β€” quarter-turned photos previously scaled boxes/polygons by the swapped dimensions and produced mismatched mask shapes. The OpenCV-free fallback backend's imread/imdecode now apply EXIF orientation too, matching OpenCV's behavior for every read except IMREAD_UNCHANGED β€” previously the same file loaded with a different shape depending on whether opencv-python was installed. (#2577)
  • sv.DetectionDataset.from_pascal_voc no longer fails on annotations with decimal coordinates (e.g. <xmin>48.5</xmin>), which aborted the whole load with ValueError: invalid literal for int() β€” Datumaro, which CVAT uses for its exports, writes VOC this way. Box coordinates are now read as floats and keep their precision; polygon vertices are rounded after the 1-index offset, as the YOLO and LabelMe loaders already do; non-finite values are still rejected. (#2568)
  • sv.DetectionDataset.from_pascal_voc no longer skips .bmp, .tif, .tiff, and .webp images without a warning β€” the loader only listed .jpg/.jpeg/.png, even though from_yolo/from_folder_structure load those formats and as_pascal_voc writes annotations for them, so a dataset exported to Pascal VOC and read back came back smaller than it went out. It now accepts the same extensions as sv.ClassificationDataset.from_folder_structure. (#2569)

Model connectors

  • sv.Detections.from_transformers now loads Transformers v5 return_binary_maps=True instance results (a (num_instances, H, W) stack), which the v5 path previously compared against each segment's id as if it were an id-map, producing a 4-D array that mask_to_xyxy rejected with `ValueError: too many v...
Read more

supervision-0.30.3

Choose a tag to compare

@Borda Borda released this 14 Sep 13:08

0.30.3: Pose, VLM, and video/CSV crash and correctness fixes

supervision 0.30.3 is a bug-fix release closing crash and silent-correctness gaps across pose estimation, VLM parsing, video/CSV output, and geometry utilities. Non-finite key points β€” how pose estimators report an undetected joint β€” no longer produce duplicate poses that survive sv.KeyPoints.with_nms, or crash the key point annotators outright. sv.Detections.from_vlm now orders backwards box corners, closing a bug where such a box scored a false 0.0 IoU and both survived NMS as a duplicate and counted as a total miss in mAP. sv.TraceAnnotator and sv.CSVSink no longer crash or silently drop columns on the first frame with no detections β€” a case every non-ByteTrack tracker pipeline hits. sv.process_video no longer hangs forever when max_frames exceeds the video length. Continuing 0.30.2's numeric-correctness theme, sv.pad_boxes and sv.scale_boxes are fixed against integer overflow. No breaking API changes, no new public API.

✨ Spotlights / highlights

Non-finite key points no longer produce duplicate poses or crash annotators

sv.KeyPoints.with_nms tested key point validity with xy == 0 alone, and NaN β€” how pose estimators report an undetected joint β€” is not 0. The stale joint stayed in the NMS box, so a duplicate skeleton scored False on every IoU comparison against it and survived suppression. The key point annotators (sv.VertexAnnotator, sv.EdgeAnnotator, sv.VertexLabelAnnotator, the sv.VertexEllipse*Annotator family) had the matching crash: a single undetected joint raised ValueError: cannot convert float NaN to integer for the whole frame. Both now skip non-finite coordinates, matching sv.KeyPoints.as_detections.

keypoints = sv.KeyPoints(xy=xy, confidence=confidence)
keypoints.with_nms(
    threshold=0.5
)  # duplicate skeletons with a NaN joint are now suppressed

sv.Detections.from_vlm no longer scores a false IoU miss on backwards box corners

A VLM that emits a corner pair backwards produced an xyxy row with x_min > x_max. Nothing downstream caught it: sv.box_iou_batch clamps intersection width at zero, so the box scored 0.0 IoU against itself β€” surviving NMS as a duplicate and counting as a total miss in mAP β€” while box_area still reported a plausible positive value. Every VLM parser now orders each box's corners before returning it.

sv.TraceAnnotator and sv.CSVSink no longer crash or silently corrupt output on an empty-detections frame

sv.TraceAnnotator.annotate raised ValueError: The tracker_id field is missing on the first frame with no detections, for every tracker except sv.ByteTrack. Such a frame now draws nothing and still advances the frame counter, so trace_length stays a window over elapsed frames rather than only over populated ones. sv.CSVSink had a quieter failure: an empty batch fixed the CSV header without the data/custom_data columns, and every later row was silently truncated to that schema β€” dropping fields like class_name for the whole file. The header is now fixed by the first batch that actually carries detections.

sv.process_video no longer hangs forever when max_frames exceeds the video length

The reader thread failed on the out-of-range end before enqueuing its sentinel, leaving the main loop blocked on the read queue indefinitely. max_frames is now capped at the video length, and any reader-thread error surfaces as RuntimeError("Reader thread raised: ...") instead of stalling the call.

Integer-coordinate overflow fixed in sv.pad_boxes and sv.scale_boxes

Both computed intermediate values that could overflow or silently wrap for large integer coordinates (e.g. large int32/uint16/int64 boxes). Both now use overflow-safe arithmetic.

xyxy = np.array([[10, 20, 30, 40]], dtype=np.int64)
sv.pad_boxes(xyxy=xyxy, px=5, py=10)  # int64 output, no wraparound

pad_boxes changes return dtype for integer input β€” see the migration guide below.

sv.scale_image and sv.resize_image(keep_aspect_ratio=True) no longer crash on an extreme aspect ratio or tiny scale factor

A small enough factor β€” or an aspect ratio too extreme for the target box β€” could round an output axis down to 0, and cv2.resize raised an assertion naming nothing the caller passed. Each axis now keeps at least one pixel. Two callers inherit the fix: sv.letterbox_image could not fill the resolution it was asked for, and sv.CropAnnotator with scale_factor < 1 aborted the whole frame as soon as one detection box was a few pixels across.

πŸ”„ Migration guide

No breaking API changes. One fix changes return dtype for integer input:

  • sv.pad_boxes β€” integer xyxy now returns int64 (or float64 if a padded coordinate exceeds the int64 range), instead of the input's original integer dtype, which could silently overflow or wrap for small dtypes like int16/uint8.

If your code assumes pad_boxes preserves the input's exact dtype (e.g. reusing the result as an int16 array), cast explicitly: sv.pad_boxes(...).astype(np.int16).

sv.scale_boxes also fixes an integer-overflow bug, but its return dtype was already float64 for integer input before this release β€” unaffected.

πŸ“ Notable changes

πŸ”§ Fixed

  • sv.Detections.from_ultralytics now assigns the placeholder class ID 0 to every mask in a masks-only result, instead of sequential IDs across masks that belong to the same image. (#2566)
  • sv.pad_boxes now computes integer-coordinate padding without overflow or unsigned casting errors. (#2565)
  • sv.scale_image and sv.resize_image(keep_aspect_ratio=True) no longer derive a zero-sized target, fixing a crash reached via sv.letterbox_image and sv.CropAnnotator. (#2564)
  • sv.KeyPoints.with_nms no longer stops suppressing duplicate skeletons as soon as a key point is non-finite. (#2563)
  • sv.tint_image no longer tints the caller's own image array in place. (#2562)
  • sv.LineZone no longer consumes the triggering_anchors iterable during validation, so a generator or map passed in is no longer exhausted before the first trigger() call. (#2561)
  • The key point annotators now skip key points whose coordinates are not finite instead of raising ValueError. (#2560)
  • sv.PolygonZone now rejects a polygon with fewer than three vertices instead of building a zone that can never trigger; sv.Detections.from_vlm now orders each parsed box's corners. (#2554)
  • sv.ClassificationDataset.as_folder_structure now rejects images that would overwrite the same class-relative filename before writing any files. (#2551)
  • sv.process_video no longer hangs forever when max_frames is larger than the number of frames in the video. (#2546)
  • sv.filter_polygons_by_area and sv.approximate_polygon now preserve local geometry for large-origin integer and float64 polygons. (#2542)
  • sv.TraceAnnotator.annotate no longer raises on an empty-detections frame; sv.CSVSink no longer lets an empty batch fix the CSV header. (#2539)
  • sv.scale_boxes now preserves exact integer intermediates, preventing overflow and scaled-corner rounding errors for large integer-coordinate boxes. (#2541)
  • A release's own version-pinned docs no longer show the outdated-version banner on the day it ships, and the docs-publish and canonical-backfill workflows now share one gh-pages write lock instead of racing each other. (#2536)

πŸ† Contributors

  • kevin (@kevin9327) β€” fixed KeyPoints.with_nms/key point annotators crashing on non-finite key points, tint_image image aliasing, LineZone generator exhaustion, and the scale_image/resize_image zero-target crash
  • Durgamani Sasikumar (@tedo001, LinkedIn) β€” fixed PolygonZone/from_vlm box-corner ordering and the TraceAnnotator/CSVSink empty-frame crash
  • S B Pranay (@pranaysb, LinkedIn) β€” fixed integer overflow in scale_boxes and dtype loss in filter_polygons_by_area/approximate_polygon
  • JiantaoPeng (@PengJianT) β€” fixed from_ultralytics masks-only class ID sizing
  • trueoneplusone (@trueoneplusone) β€” fixed integer overflow in pad_boxes
  • Andrew Barnes (@Bortlesboat, LinkedIn) β€” fixed classification export filename collisions
  • Abhijith Neil Abraham (@abhijithneilabraham, LinkedIn) β€” fixed process_video hanging when max_frames exceeds the video length
  • Jirka Borovec (@Borda, LinkedIn) β€” fixed the release-day outdated-docs banner

Full changelog: 0.30.2...0.30.3

supervision-0.30.2

Choose a tag to compare

@Borda Borda released this 04 Sep 07:22

0.30.2: Detection numeric-correctness fixes

supervision 0.30.2 fixes three silent numeric-correctness bugs in the detection utilities β€” integer box areas that could wrap negative on large boxes, and two coordinate converters that truncated fractional values on integer input β€” plus an InferenceSlicer determinism fix that restores its documented source-order result guarantee under multithreading. A set of versioned-docs reliability fixes rounds out the release. No breaking API changes, no new public API.

✨ Spotlights / highlights

sv.Detections.box_area no longer overflows to a negative number

Integer-coordinate box area now computes in float64. A large int32 box (50000 x 50000) previously wrapped to a negative area.

detections = sv.Detections(xyxy=np.array([[0, 0, 50000, 50000]], dtype=np.int32))
detections.box_area  # array([2.5e+09]) β€” was negative before the fix

xcycwh_to_xyxy / denormalize_boxes stop truncating integer boxes

Both converters wrote fractional half-extent or scaled coordinates into a copy of the integer input, silently truncating toward zero β€” relevant when converting quantized VLM output (e.g. boxes on a 0..1000 grid).

xcycwh_to_xyxy(np.array([[10, 10, 5, 5]], dtype=np.int32))
# array([[ 7.5,  7.5, 12.5, 12.5]]) β€” the fractional coordinate 7.5 is no longer truncated to 7

InferenceSlicer merges results in source order under multithreading

Slice results now merge in source order under thread_workers > 1, restoring the ordering guarantee its docstring documents. Row order β€” and, for tied confidences, which overlapping box survives with_nms/with_nmm β€” no longer varies between runs on identical input.

πŸ”„ Migration guide

No breaking API changes, but three fixes above change return dtype for integer input:

  • sv.Detections.box_area / .area β€” integer xyxy now returns float64 (was the input's integer dtype, which could silently overflow)
  • sv.xcycwh_to_xyxy β€” integer input now returns float64 (was truncated integer output)
  • sv.denormalize_boxes β€” integer input now returns float64 (was truncated integer output)

If your code indexes arrays with these outputs (e.g. image[y1:y2, x1:x2]), a float64 result raises TypeError: slice indices must be integers. Cast explicitly where integer indices are required: xcycwh_to_xyxy(boxes).astype(int).

πŸ“ Notable changes

πŸ”§ Fixed

  • sv.Detections.box_area (and sv.Detections.area for axis-aligned boxes) now computes integer-coordinate box areas in float64, preventing integer overflow for large boxes. (#2514)
  • sv.xcycwh_to_xyxy no longer truncates coordinates for integer input arrays. (#2515)
  • sv.denormalize_boxes no longer truncates coordinates for integer input arrays. (#2516)
  • sv.InferenceSlicer now merges slice results in source order when thread_workers > 1, restoring the ordering guarantee its docstring documents. (#2517)
  • Docs deployment for latest no longer fails with error: version 'latest' already exists when latest exists as an alias of a released version. (#2512, #2513)
  • Versioned documentation builds now emit a valid /latest/search/ SearchAction URL when Mike removes the trailing slash from site_url; docs CI renders the custom theme under Mike version contexts to protect the URL, version banners, and star JSON-LD. Applies to future builds going forward. (#2529)
  • Versioned documentation banners now adjust MkDocs Material's desktop sidebar inline layout and scroll height without shifting the mobile navigation drawer. (#2532)
  • Versioned documentation deploys now export the version being built, so the outdated-version banner reaches readers of the develop tree; the workflow also now backs up the pre-rewrite gh-pages tip to a timestamped branch before committing over it. Applies to future builds going forward. (#2533)
  • The canonical-backfill workflow now reports (in the job summary) rewritten canonicals whose target page does not exist under latest/, and backfills the outdated-version banner itself into already-published archive trees, patching the empty banner markup those pages already carry rather than rebuilding them. (#2534)

πŸ† Contributors

  • Advait Shukla (@AdvaitS) β€” fixed integer overflow in box_area
  • Guillaume Flambard (@guillaume-flambard, LinkedIn) β€” fixed integer truncation in denormalize_boxes and xcycwh_to_xyxy
  • Roshan Sharma (@roshaninfordham, LinkedIn) β€” fixed InferenceSlicer result ordering under multithreading
  • Jirka Borovec (@Borda, LinkedIn) β€” versioned-docs banner layout, MIKE_DOCS_VERSION export, gh-pages backup, and canonical-backfill fixes (#2529, #2532, #2533, #2534)

Full changelog: 0.30.1...0.30.2

supervision-0.30.1

Choose a tag to compare

@Borda Borda released this 24 Aug 23:13

0.30.1: Numeric-precision and stability fixes

supervision 0.30.1 is a bug-fix patch release. It corrects numeric-precision issues that only surface on specific inputs β€” large-coordinate oriented boxes (geospatial data, stitched frames), large integer boxes for box_iou, and rotated tracks in DetectionsSmoother β€” where prior versions could silently return imprecise or self-inconsistent results instead of erroring. It also fixes a duplicate-libavdevice-load crash risk on macOS when both av and opencv-python are installed, plus smaller fixes to list_files_with_extensions and the cv2-free RGBA fallback. No public API was added or removed, and no signature changed β€” a drop-in upgrade from 0.30.0 for virtually all users. See Migration guide below for the one narrow exception (box_iou on complex-valued coordinates) and for the precision caveats on the numeric fixes.

✨ Spotlights / highlights

1. Oriented-box area/IoU precision fix for large coordinates

sv.Detections.area and sv.oriented_box_iou_batch now translate OBB coordinates to a local origin before floating-point math. Previously, large-coordinate inputs could lose enough precision that a box's IoU with itself collapsed below 1.0.

pair_origin = np.minimum(origin_i, origin_j)
offset_i = (origin_i - pair_origin).astype(np.float32, copy=False)
offset_j = (origin_j - pair_origin).astype(np.float32, copy=False)

2. sv.box_iou no longer overflows on large integer boxes

Area computation now takes coordinate differences before casting to float, avoiding int32 overflow. For realistic coordinate magnitudes, box_iou's scalar result now matches box_iou_batch.

3. Duplicate libavdevice crash fixed on macOS

import supervision no longer loads PyAV's native libraries when the OpenCV backend is active β€” PyAV is now imported lazily, only where it's used, preventing a duplicate libavdevice warning (and possible crash) when both av and opencv-python are installed.

4. DetectionsSmoother keeps oriented-box corners consistent

Smoothed OBB corners are now aligned (start index + winding) to a reference before averaging, so rotated tracks smooth correctly instead of averaging mismatched corner orderings.

5. sv.get_polygon_center precision fix for large-coordinate polygons

Centroid calculation now translates to the first vertex and computes in float64 before adding the origin back, preventing integer overflow and precision loss for realistic coordinate magnitudes.

πŸ”„ Migration guide

No public signature changed. One item below (box_iou on complex coordinates) does make one specific previously-succeeding call now raise β€” narrow and deliberate, not classified as breaking since complex-valued box coordinates were never a documented/supported input. The rest only change output values for inputs that were already edge cases:

  • sv.box_iou on complex-valued coordinates: previously silently discarded the imaginary part and returned a real number. Now raises TypeError("box coordinates must be real-valued").
  • OBB precision fixes: results for oriented boxes with large coordinates or rotated tracks may differ slightly from 0.30.0 β€” the new values are the corrected ones. Re-calibrate any hardcoded IoU/area thresholds tuned against the old (imprecise) output.
  • sv.box_iou / box_iou_batch agreement: for realistic integer coordinate magnitudes (below 2^53), box_iou's scalar result now matches box_iou_batch. Not a universal guarantee β€” box_iou subtracts before casting to float, box_iou_batch still casts to float64 before subtracting, so the two can diverge at coordinates β‰₯ 2^53 (~9 quadrillion), far outside any real use case.

πŸ“ Notable changes

πŸš€ Added

  • RF-DETR example scripts (rfdetr_example.py) added to the count_people_in_zone, heatmap_and_track, speed_estimation, tracking, and traffic_analysis bundled examples. (#2497)

🌱 Changed

  • sv.box_iou now raises TypeError for complex-valued box coordinates instead of silently discarding the imaginary part. (#2485)
  • Performance: DetectionsSmoother.update_with_detections now checks active tracker IDs via set membership instead of scanning per tracked object. No output changes. (#2496)
  • Documentation and API-reference examples now default to RF-DETR instead of Ultralytics YOLO. (#2493, #2494, #2497)

πŸ”§ Fixed

  • RF-DETR speed estimation now measures elapsed source-frame intervals, including gaps when tracked detections are temporarily missed. (#2497)
  • sv.get_polygon_center now calculates polygon centroids in translated float64 coordinates, preventing integer overflow and precision loss for realistic-magnitude large-coordinate polygons. (#2491)
  • sv.Detections.area and sv.oriented_box_iou_batch now translate oriented-box coordinates to local origins before floating-point math, preventing self-IoU collapse for large-coordinate inputs. (#2492)
  • DetectionsSmoother now keeps oriented-box corners aligned with smoothed xyxy geometry, including rotated tracks and mixed metadata windows. (#2489)
  • sv.box_iou now calculates overlap in float64, preventing int32 area overflow for large boxes; its scalar result now matches sv.box_iou_batch for realistic coordinate magnitudes. (#2485)
  • sv.list_files_with_extensions no longer includes directories when listing all files without an extension filter. (#2486)
  • sv.pillow_to_cv2 now accepts RGBA images when the cv2-free fallback backend is active, matching OpenCV by dropping alpha and returning BGR channels. (#2488)
  • import supervision no longer loads PyAV's native libraries when the OpenCV backend is selected; PyAV is now imported lazily on first use, preventing a duplicate libavdevice warning (and possible crash) on macOS when both av and opencv-python are installed. (#2509)
  • Docstring examples converted to executed doctests across annotators/core.py, dataset/formats/coco.py, dataset/formats/createml.py, and Detections.from_vlm; previously wrong documented outputs corrected for several VLM examples. (#2474, #2475, #2479, #2484)

Also in this release: routine dependency bumps (dependabot: astral-sh/setup-uv, wheel, pymdown-extensions x2, pypa/gh-action-pypi-publish, twine, cryptography), CI/docs-workflow maintenance, and test-only additions (geometry contract test, sklearn parity test) β€” none change installed package behavior. (#2028, #2470, #2472, #2473, #2480, #2481, #2482, #2483, #2499, #2501, #2506, #2507, #2508)

πŸ† Contributors

  • lawliet (@lawliet206) β€” oriented-box area/IoU precision fix at large coordinate origins
  • Zhewen Tan (@tandede) β€” polygon centroid overflow/precision fix
  • Tamil Adhavan S K (@adhavan18, LinkedIn) β€” DetectionsSmoother oriented-box corner alignment fix; added geometry contract test
  • NIKHIL (@Nikhi00718) β€” fixed box_iou int32 overflow; added complex-coordinate TypeError guard
  • shao (@shaoming11, LinkedIn) β€” DetectionsSmoother tracker-ID lookup performance improvement
  • Tyyyy (@uczltw6) β€” RGBA image support in the cv2-free fallback conversion
  • ZZZZZ (@BruceWae) β€” fixed list_files_with_extensions to exclude directories
  • FootysHands (@ayo0la) β€” converted docstring examples to doctests; corrected wrong documented VLM example outputs
  • Swapnil Gautam (@Swapnil-gautam) β€” converted docstring examples to doctests in dataset format modules
  • Daniiiil1 (@Daniiiil1) β€” added sklearn parity test for metrics
  • Christoph Deil (@cdeil, LinkedIn) β€” repo maintenance docs

Full changelog: 0.30.0...0.30.1

supervision-0.30.0

Choose a tag to compare

@Borda Borda released this 04 Aug 17:36
813c429

v0.30.0: Run supervision without OpenCV

supervision 0.30.0 makes OpenCV optional. A new private _cv2/ backend (NumPy and Pillow, with PyAV for the video path) reimplements every OpenCV call the library needs, so supervision now runs on opencv-python-headless β€” or no OpenCV wheel at all β€” instead of crashing on import. This release also adds Soft-NMS, LabelMe and CreateML dataset formats, GeoTIFF-aware windowed reads for InferenceSlicer, and ships five breaking changes, most notably OpenCV no longer being installed by default, JSONSink switching to native JSON types, and mask_non_max_merge computing exact mask overlap instead of a downscaled approximation. Python 3.9 support is dropped β€” 3.10 is now the minimum.

✨ Spotlights / highlights

Run supervision without OpenCV

import supervision as sv

window = sv.ImageWindow("frame")
for frame in sv.get_video_frames_generator("input.mp4"):
    window.show(frame)
    if window.wait_key(1) == "q":
        break

The largest change in this release: OpenCV stays the default backend when installed, but supervision no longer requires it β€” there's no opencv-python extra anymore either. sv.ImageWindow replaces cv2.imshow/cv2.waitKey for display. av>=14.2 is now a required dependency for the PyAV video path during this transition. See the OpenCV migration guide.

Soft-NMS

detections = sv.Detections.from_ultralytics(result)
softened = detections.with_soft_nms(sigma=0.5)
filtered = detections.with_soft_nms(sigma=0.5, score_threshold=0.3)

sv.Detections.with_soft_nms (plus sv.box_soft_non_max_suppression / sv.mask_soft_non_max_suppression) rescales overlapping detections' confidence instead of discarding them outright β€” useful in crowded scenes where hard NMS drops valid overlapping objects.

New dataset formats + GeoTIFF-aware, batched slicing

dataset = sv.DetectionDataset.from_labelme(
    images_directory_path="images/",
    annotations_directory_path="annotations/",
)

import rasterio

with rasterio.open("RGB.byte.tif") as raster:
    slicer = sv.InferenceSlicer(callback=my_model_callback, batch_size=4)
    detections = slicer(raster)

DetectionDataset.from_labelme/as_labelme and from_createml/as_createml join the existing COCO/YOLO/Pascal-VOC converters. sv.InferenceSlicer can now read an open rasterio dataset window-by-window for multi-GB aerial/drone GeoTIFFs without loading the whole image (pip install "supervision[geotiff]"), and accepts batch_size for batched-callback inference.

sv.load_image_from_url

image = sv.load_image_from_url("https://media.roboflow.com/notebooks/examples/dog.jpeg")

Load an image straight from an HTTP(S) URL as an OpenCV array, with optional on-disk caching.

πŸ”„ Migration guide

Five breaking changes. Most require no code changes beyond a type check or threshold recalibration. The two that need action from most users: the OpenCV install change below, and the Python 3.10 floor.

OpenCV is no longer installed by default. If a compatible cv2 is already importable in your environment, nothing changes for you β€” it's still preferred automatically. Otherwise install one wheel family yourself (pip install opencv-python or opencv-python-headless) if you need OpenCV-specific behavior, then restart the process β€” cv2 is detected once at import time. sv.ImageWindow replaces cv2.imshow/cv2.waitKey. Full guide: docs/how_to/opencv_migration.md.

Python 3.10+ is now required β€” 3.9 reached end-of-life in October 2025.

sv.JSONSink now emits native JSON types, not strings:

# before 0.30.0
row["score"] == "0.85"  # str
row["is_valid"] == "True"  # str

# after 0.30.0
row["score"] == 0.85  # float
row["is_valid"] is True  # bool

sv.CSVSink stays textual, but its per-row custom-data slicing now matches JSONSink.

sv.mask_non_max_merge computes exact mask overlap, not a downscaled approximation, and ignores the now-deprecated mask_dimension parameter (kept for signature compatibility, removal in 0.33.0). Re-tune your overlap threshold after upgrading. Passing overlap_metric/mask_dimension positionally still works β€” the values are still honored β€” but now emits a DeprecationWarning; pass them by keyword to silence it. More than five positional arguments raises TypeError.

Detections.merge() on mixed dense + CompactMask inputs now returns a CompactMask, not a plain ndarray:

merged = sv.Detections.merge([dense_detections, compact_mask_detections])
isinstance(merged.mask, np.ndarray)  # was True, now False β€” it's a CompactMask

Only affects code that explicitly merges a CompactMask-carrying Detections object with a dense-mask one yourself β€” InferenceSlicer, DetectionsSmoother, and with_nms/with_nmm always merge type-homogeneous lists internally, so they're unaffected. The all-dense merge path is also unchanged. This is a substantial performance win: ~2500Γ— less peak memory, ~13Γ— faster on a 1080p frame with 40 detections. If you need the old return type without touching every call site: call merged.mask = merged.mask.to_dense() right after merge(), or avoid producing CompactMask in the first place (Detections.from_inference(compact_masks=False), the default).

supervision also now requires av>=14.2 as an install-time dependency for the PyAV cv2-free video path β€” this doesn't change any API, so it isn't counted as breaking, but pinned/vendored environments should account for it.

Deprecation removals pushed back one release: ByteTrack, supervision.keypoint, normalized_xyxy, and supervision.dataset.utils RLE compatibility shims β€” originally scheduled for removal in 0.30.0 β€” are now scheduled for 0.31.0 instead, giving a full transition window.

πŸ“ Notable changes

πŸš€ Added

  • sv.load_image_from_url β€” load an HTTP(S) image as an OpenCV array, with optional on-disk caching (#2372)
  • cv2-free PyAV video fallback + private _cv2 backend facade β€” image/geometry/drawing/text/video without OpenCV (#2430, #2431, #2432, #2433, #2435, #2438, #2439, #2440, #2441, #2443)
  • sv.ImageWindow β€” tkinter+Pillow desktop window replacing cv2.imshow/cv2.waitKey (#2320)
  • Soft-NMS β€” sv.box_soft_non_max_suppression, sv.mask_soft_non_max_suppression, sv.Detections.with_soft_nms (#1624)
  • sv.VLM.GOOGLE_GEMINI_3_5 β€” Detections.from_vlm parses Gemini 3.5 output (#2449)
  • get_video_frames_generator(prefetch=...) β€” background-thread decode into a bounded queue (#2273)
  • PolygonZone(require_all_anchors=...) β€” toggle all-anchors vs. any-anchor containment (#2272)
  • KeyPoints.merge() β€” combine a list of KeyPoints, mirroring Detections.merge (#2412)
  • BaseAnnotator.requires_mask β€” class-level flag on all annotators (#2370)
  • CompactMask.from_coco_rle + Detections.from_inference(compact_masks=True) (#2367)
  • CompactMask.image_shape property (#2383)
  • sv.mask_to_roi β€” exclusive mask-bound helper for slicing/crops (#2416)
  • DetectionDataset.from_labelme/as_labelme (#2299)
  • DetectionDataset.from_createml/as_createml (#2284)
  • InferenceSlicer GeoTIFF support β€” sv.WindowedRasterDataset, pip install "supervision[geotiff]" (#2281)
  • InferenceSlicer(batch_size=...) β€” batched callback contract (#1239)
  • ConfusionMatrix.benchmark(save_directory_path=...) β€” adaptive TP/FP/FN validation-mosaic export (#2271)
  • HeatMapAnnotator.reset(), TraceAnnotator.reset(), DetectionsSmoother.reset() β€” clear accumulated per-stream state, so a single instance can be reused across independent streams (#2418)
  • AREA_DATA_FIELD config constant (#2428)
  • sv.denormalize_boxes and sv.xyxyxyxy_to_xyxy now exported at the top level

⚠️ Breaking Changes

  • OpenCV no longer installed by default; no OpenCV extra (#2443)
  • Python 3.10+ required β€” 3.9 dropped (#2260, #2381)
  • sv.JSONSink emits native JSON types instead of strings; sv.CSVSink custom-data slicing now matches JSONSink (#2400)
  • sv.mask_non_max_merge computes exact overlap, ignores mask_dimension, positional overlap_metric/mask_dimension deprecated (#2400)
  • Detections.merge() on mixed dense + CompactMask inputs returns CompactMask (#2383)

🌱 Changed

  • DetectionDataset/ClassificationDataset equality now compares ordered classes lists, not an unordered set
  • supervision now requires av>=14.2 as an install-time dependency for the cv2-free video fallback β€” no API change (#2438)
  • Deprecation-window delays: ByteTrack, supervision.keypoint, normalized_xyxy, dataset-utils RLE compat removals moved 0.30.0 β†’ 0.31.0
  • Perf: count_nonzero mask pixel counts (#2361), vectorized box_iou_batch_with_jaccard (#2359), faster mask-annotation ROI blending (#2368), fewer corner circles on square label backgrounds (#2346), less compact-mask materialization in the polygon annotator (#2369)
  • Geometry-aware IoU/area dispatch centralized (#2374)

πŸ”§ Fixed

  • sv.Recall tracks prediction-only classes, matching Precision/F1Score (#2467, #2468)
  • DetectionDataset.from_pascal_voc no longer raises on background images, with or without force_masks=True (#2463, #2469)
  • import supervision no longer surfaces the deprecated ByteTrack warning
  • Reopening sv.CSVSink/sv.JSONSink starts a fresh session β€” no stale rows or header (#2459)
  • from_vlm Gemini 2.0/2.5/3.5 salvages valid entries from partially malformed JSON arrays (#2449)
  • save_coco_annotations/as_coco read image sizes from headers, no pixel decode for labels-only export (#2442)
  • sv.F1Score no longer emits a spurious div-by-zero `RuntimeWarni...
Read more

supervision-0.29.1

Choose a tag to compare

@Borda Borda released this 23 Jun 19:54

What's new

πŸš€ KeyPoints.with_nms() β€” NMS for pose estimation

import supervision as sv

key_points = model.predict(image)  # sv.KeyPoints
key_points = key_points.with_nms(threshold=0.5)  # removes duplicate skeletons

Derives axis-aligned bounding boxes from each skeleton's valid (non-zero and visible) keypoints, then applies standard box NMS. Supports class_agnostic mode and any OverlapMetric (IOU, IOS). Raises ValueError if detection_confidence is not set.

feat-keypoints-with-nms-compressed.mp4

Notable changes

Bug fixes

  • sv.DetectionDataset.as_pascal_voc no longer mutates bounding boxes (#2341) Previously, every export shifted every bounding box by +1 px in-place. A second call compounded the shift. Fixed by rebinding to a new array; on-disk XML output is unchanged.

  • sv.Precision and sv.F1Score correctly count background false positives (#2331) Predictions on images with no ground-truth objects, and predictions of classes absent from any annotation, were previously ignored. Under MICRO and MACRO averaging they are now counted as false positives. WEIGHTED averaging is unchanged. Users should re-evaluate existing metric results after upgrading.

  • sv.DetectionsSmoother works with confidence-free detections (#2333) The smoother no longer raises when detections have no confidence scores. Confidence is averaged over the frames that carry it; tracks without any confidence produce None.

  • sv.Detections.from_vlm is robust to malformed Gemini/Qwen output (#2342) Valid JSON that is not a list, or whose elements are not dicts, now degrades to empty Detections instead of raising TypeError. A malformed mask value in Gemini 2.5 responses no longer misaligns the xyxy/confidence/masks arrays.

  • sv.JSONSink serializes NumPy scalars in custom_data (#2334) np.int64 frame indices and other NumPy scalars in custom_data no longer raise TypeError at flush time. NumPy arrays are serialized as lists. The file handle closes even when serialization fails.

  • sv.approximate_polygon respects the point-count budget (#2332) The function now returns at most floor(N * (1 - percentage)) points (minimum 3). Previously it could return more points than requested. epsilon_step is now validated to be positive.

  • COCO export preserves all segments for multi-part masks (#2322) Previously, only the first polygon was written when a non-crowd detection had disjoint mask segments. All polygon parts are now written.

Performance

  • sv.HaloAnnotator is ~4Γ— faster with CompactMask detections (#2339) HaloAnnotator now uses the same optimized CompactMask paint path as MaskAnnotator. Previously it materialized each mask full-frame; now it operates on the bounding-box crop. Annotated output is unchanged.

  • Mask IoU uses less peak memory (#2323) Mask IoU computation now uses matrix multiplication on flattened masks instead of an explicit (N, M, H, W) tensor. For masks larger than 4096Γ—4096 px, computation promotes to float64 automatically. Results are numerically identical.

  • sv.mask_to_xyxy and sv.KeyPoints.as_detections vectorized (#2330) Both functions now use batched NumPy operations instead of per-element loops. Outputs are bit-identical.


Contributors

  • Ruben Haisma (@RubenHaisma, LinkedIn) β€” VLM robustness, Pascal VOC export fix, DetectionsSmoother, JSONSink, metrics correctness, polygon budgeting, vectorization
  • Agis Kounelis (@kounelisagis, LinkedIn) β€” HaloAnnotator perf, mask IoU matmul, mask_to_xyxy/KeyPoints.as_detections vectorization, OBB cookbook
  • Piotr Skalski (@SkalskiP, LinkedIn) β€” KeyPoints.with_nms()
  • Abdelrahman Gomaa (@abdogomaa201099, LinkedIn) β€” COCO multi-polygon export

Full Changelog: 0.29.0...0.29.1