Repository navigation
Releases: roboflow/supervision
Release list
supervision-0.30.8
v0.30.8 β Sharper video, cleaner labels
Video, YOLO labels, masks, VLM parsing and mAP all get more accurate.
VideoSinkkeeps OpenCV's video quality when OpenCV isn't installed.from_yoloreads pose labels as boxes instead of polygons.MeanAveragePrecisionscores class-agnostic runs right when only one side has class IDs.from_vlmreturns one Florence-2 detection per object, not one per polygon.from_inferencemasks no longer drift up to a pixel up and left.
Drop-in upgrade. Without OpenCV, videos get larger; YOLO labels with a negative width or height now raise ValueError.
β¨ Spotlights / highlights
sv.VideoSink and sv.process_video keep quality without OpenCV
The PyAV fallback left the encoder bit rate unset, so mp4v and MJPG files came out at under half of what cv2.VideoWriter writes. It now uses OpenCV's rate settings for every codec except H.264. Files get larger and vp09 may encode more slowly; codec="avc1" keeps files small where an H.264 encoder is available. (#2661)
import supervision as sv
video_info = sv.VideoInfo.from_video_path("in.mp4")
with sv.VideoSink("out.mp4", video_info) as sink: # OpenCV not installed
for frame in sv.get_video_frames_generator("in.mp4"):
sink.write_frame(frame)
# before: mp4v written at under half OpenCV's bit rate, visibly softer
# now: same bit rate OpenCV's writer usesYOLO pose labels load as boxes
A pose row is a box followed by keypoints. from_yolo used to parse the whole row as a polygon, giving wrong boxes and masks nobody asked for. It now reads the box and skips the keypoints that kpt_shape declares. (#2655)
ds = sv.DetectionDataset.from_yolo(
images_directory_path="pose/images",
annotations_directory_path="pose/labels",
data_yaml_path="pose/data.yaml", # kpt_shape: [17, 3]
)
# before: polygon-parsed boxes and masks
# now: one box per rowClass-agnostic mAP with one-sided class IDs
A perfect match scored zero when only one side carried class IDs, such as SAM proposals checked against labeled ground truth. With class_agnostic=True, both sides now count as one class.
One Florence-2 detection per object (#2648)
Florence-2 returns a segmented object as a list of polygons, one per connected region. An object split in two used to come back as two detections; the polygons now merge into one mask with one box around all of them.
Roboflow masks sit on the right pixels (#2649)
from_inference truncated sub-pixel polygon vertices, shifting each mask up and left by up to a pixel. Vertices are now rounded, the way the COCO, YOLO, LabelMe and Pascal VOC loaders already do.
π Migration guide
No migration required for this release.
π Notable changes
π± Changed
sv.DetectionDataset.from_yoloraisesValueErrornaming the annotation file when a label has a negative width or height; it used to load a box withx_minpastx_max, which madeDetections.areanegative and skewed IoU and NMS.as_yolonow orders the corners of a reversed box before measuring, so it no longer writes a file the loader refuses. (#2663)
π§ Fixed
sv.VideoSinkandsv.process_videowritemp4v,MJPGand other non-H.264 video at OpenCV's bit rate when OpenCV isn't installed. A frame rate of zero or less now raisesRuntimeErrorinsv.VideoSink, as it does with OpenCV. (#2661)sv.DetectionDataset.from_yoloreads the box of Ultralytics pose labels and skips their keypoints; akpt_shapeother than[K, 2]or[K, 3]raisesValueError. (#2655)sv.DetectionDataset.from_yolonames a malformed annotation line, one with too few values or, for OBB, not nine, in aValueErrorinstead of failing on an array shape. (#2665)sv.metrics.MeanAveragePrecision(class_agnostic=True)treats detections without class IDs as the same class as labeled ones, and unsigned class ID arrays no longer raiseOverflowErroron NumPy 2. (#2650)sv.Detections.from_vlmwithsv.VLM.FLORENCE_2merges the polygons of one instance into one detection for<REFERRING_EXPRESSION_SEGMENTATION>and<REGION_TO_SEGMENTATION>, and skips instances with no usable polygon. (#2648)sv.Detections.from_vlmwithsv.VLM.QWEN_2_5_VLorsv.VLM.QWEN_3_VLrecovers complete detections from a response cut off inside abbox_2darray or right after a complete object. (#2666)sv.Detections.from_inferencerounds polygon vertices to the nearest pixel before rasterising masks, and raisesValueErrorfor NaN or infinite vertices. (#2649)sv.xyxy_to_maskreturns an empty mask for a box entirely left of or above the image when its maximum coordinate is a negative fraction. (#2646)sv.Detections.get_anchors_coordinatescomputes axis-aligned midpoint anchors without integer overflow. (#2660)sv.LineZone.triggerages crossing history on frames whose detections lacktracker_id, so a reused track ID no longer creates a false crossing after the track expired. (#2644)sv.LineZoneAnnotator(text_orient_to_line=True)no longer raisesTypeErrorwithout OpenCV for lines drawn right to left. (#2659)- The Ultralytics, Inference and YOLO-NAS speed estimation examples measure elapsed time from frame indices; a vehicle missed in one frame of three was reported about 44% too fast. (#2654)
examples/speed_estimation/rfdetr_example.pyno longer raisesAttributeError:supervision._cv2now providesgetPerspectiveTransformandperspectiveTransform. (#2652)
π Contributors
- Mohammad Hijjawi (@MohammadHijjawi97, LinkedIn) β fixed video quality without OpenCV, Florence-2 instance merging and Roboflow mask rounding.
- Kari Pikkarainen (@kari-pikkarainen, LinkedIn) β fixed speed estimation timing, the NumPy
flipfallback and the perspective-transform fallbacks. - Miral Amin (@aminmiral) β made YOLO loading reject negative extents and name malformed lines.
- NIKHIL (@Nikhi00718) β fixed
LineZonehistory expiry and anchor overflow. - kevin (@kevin9327) β made Qwen parsing recover from cut-off responses.
- A Aswanth Raj (@aswanth-07, LinkedIn) β fixed class-agnostic mAP.
- Devulapalli Naga Sri Vaishnavi (@Vaishnavi220506) β fixed masks for off-frame fractional boxes.
- JANG BYUNGKUN (@8rulerstar) β fixed YOLO pose label loading.
Automated contributions: @dependabot
Full changelog: 0.30.7...0.30.8
supervision-0.30.7
v0.30.7 β Fewer hangs, truer mAP
Eight bug fixes: video processing, dataset export, mAP and Transformers loading no longer fail or hang silently.
process_videoraises on a failed write instead of hanging.MeanAveragePrecisionmatchespycocotoolswhere recall lands on a threshold.- Dataset export can write back into the folder the images came from.
from_transformersreads semantic output that includes per-pixel scores.
Drop-in upgrade, no code changes. Stored mAP baselines may shift slightly.
β¨ Spotlights / highlights
sv.process_video stops hanging (#2636)
A frame the writer rejects, or a callback that returns None, now raises instead of blocking forever.
import supervision as sv
def callback(frame, index):
sv.BoxAnnotator().annotate(frame.copy(), detections) # forgot `return`
sv.process_video("in.mp4", "out.mp4", callback)
# before: hangs after `writer_buffer` frames
# now: TypeError naming the framesv.metrics.MeanAveragePrecision matches pycocotools (#2638)
Recall thresholds and recall are float64, as in COCOeval. A class whose recall lands exactly on a threshold used to score slightly high: mAP@50 0.7470 instead of 0.7415 in one case.
Datasets re-export in place (#2637)
ds = sv.DetectionDataset.from_coco(
images_directory_path="data",
annotations_path="data/_annotations.coco.json",
)
ds.as_coco(
images_directory_path="data",
annotations_path="data/_annotations.coco.json",
)
# before: shutil.SameFileError
# now: annotations rewritten, images left in placeTransformers semantic output with scores (#2643)
sv.Detections.from_transformers accepts results from return_segmentation_scores=True instead of raising KeyError: 'segments_info'.
π Migration guide
No migration required for this release.
π Notable changes
π§ Fixed
sv.process_videoraisesRuntimeError("Writer thread raised: ...")orTypeErrorinstead of hanging when a frame cannot be written or a callback returnsNone. (#2636)sv.metrics.MeanAveragePrecisioncomputes IoU and recall thresholds, IoUs and recall in float64, matchingpycocotools. mAP changes only for classes whose recall lands exactly on one of the 101 thresholds, by up to about 0.007 mAP@50. (#2638)sv.DetectionDataset.as_yolo,as_pascal_voc,as_coco,as_createmlandas_labelmeexport into the folder the images were loaded from instead of failing withshutil.SameFileError. (#2637)sv.Detections.from_transformersaccepts semantic segmentation output withsegmentation_scores; per-pixel scores stay out of per-detectionconfidence. (#2643)sv.crop_imageclips finite crop coordinates outside the 32-bit integer range to the image bounds, instead of wrapping to an empty crop. (#2642)sv.DetectionDataset.as_labelmegives disconnected components of one mask a shared group ID, sofrom_labelmerebuilds one detection. (#2640)sv.DetectionDataset.as_labelme,as_yoloandas_pascal_vocexport in-memory grayscale(height, width)images. (#2641)sv.KeyPoints.with_nmskeeps skeletons with zero joints instead of raising a zero-size reduction error. (#2639)
π Contributors
- kevin (@kevin9327) β stopped
process_videohangs, made mAP matchpycocotools, and fixed in-place dataset export. - Marvel Harisson (@INo-xious, LinkedIn) β fixed LabelMe grouping, grayscale dataset export and keypoint NMS.
- NIKHIL (@Nikhi00718) β fixed
crop_imageclipping and Transformers semantic output loading.
Full changelog: 0.30.6...0.30.7
supervision-0.30.6
0.30.6: Pillow-conversion, dataset-export, and line-crossing correctness fixes
supervision 0.30.6 is a patch release closing 12 library correctness bugs and landing 2 behavior refinements across datasets, annotators, key points, line-crossing counting, and the Pillow-based fallback backend, plus a new docs guide and a notebook dead-link fix. No new public API, no breaking changes. The broadest-reach fix rewrites pillow_to_cv2, the entry point every annotator except RichLabelAnnotator and several image helpers use to convert a Pillow image: it used to pass raw mode bytes straight through, so a 1-bit mask, a 16-bit depth map, an LA/PA image, or a CMYK JPEG all drew wrong or crashed. Two fixes close silent-wrong-data bugs rather than crashes: LineZone.trigger let unconfirmed tracks inflate crossing counts, and DetectionDataset.as_yolo/.as_pascal_voc silently dropped box-only objects from a mixed polygon/box COCO dataset. DetectionsSmoother no longer crashes when two tracked objects first seen on different frames carry different metadata (e.g. RF-DETR's source_image); its own docstring pipeline was broken. Every library fix ships a regression test.
β¨ Spotlights / highlights
sv.pillow_to_cv2 now converts every Pillow mode the way cv2.imread would
Raw mode bytes used to pass straight through: a 1-bit mask came back as 0/1 instead of 0/255, a 16-bit depth map wrapped modulo 256 and drew as noise, an LA/PA image crashed in cvtColor on its 2-channel array, and a CMYK JPEG drew its cyan/magenta/yellow ink as red/green/blue. This function is the entry point for every annotator except RichLabelAnnotator (which stays on the Pillow path) plus crop_image, resize_image, letterbox_image, scale_image, tint_image, grayscale_image, and plot_image whenever handed a Pillow image, so the bug reached the whole drawing surface, not just one call site. A single-channel grayscale image also drew from a read-only buffer view before this fix. Annotating into it silently failed; it now hands back a writable copy.
from PIL import Image
depth_map = Image.open("depth.tif") # mode "I;16"
scene = sv.pillow_to_cv2(depth_map)
# clipped to the 16-bit range and keeps its high byte, instead of wrapping mod 256RGB, RGBA, grayscale, and palette images are unchanged.
sv.LineZone.trigger no longer lets unconfirmed tracks inflate counts
Detections with a negative tracker_id (how trackers like ByteTrackTracker mark an unconfirmed track) all keyed under one shared id. Two distinct unconfirmed objects crossing on opposite sides in the same frame read as one track oscillating, silently inflating in_count/out_count with crossings no confirmed track ever made. Confirmed tracks (tracker_id >= 0) are counted exactly as before.
sv.DetectionDataset.as_yolo / .as_pascal_voc stop dropping box-only objects
from_coco gives an all-zero mask to any annotation without its own segmentation when a sibling annotation has one, and to every annotation when force_masks=True. The exporters only ever wrote the polygon traced from a mask, so a mixed polygon/box COCO dataset silently lost its box-only objects on export, and a box-only COCO dataset loaded with force_masks=True wrote empty YOLO label files or object-less Pascal VOC files. A detection with an empty or contour-less mask now exports as its bounding box instead.
sv.DetectionsSmoother no longer crashes on a second tracked object
Each smoothed track was built on the oldest frame in its window, so two objects first seen on different frames carried different frames' metadata (exactly what the RF-DETR / inference connectors attach), and Detections.merge rejected the mismatch. class_id and data (e.g. class_name) now follow the current frame instead of lagging up to length - 1 frames behind a class change; xyxy, confidence, and oriented-box corners are still averaged as before.
Dataset splits reject an out-of-range ratio instead of silently mis-splitting
sv.DetectionDataset.split, sv.ClassificationDataset.split (both take split_ratio), and the internal train_test_split helper they call (takes train_ratio) never validated the ratio. A finite out-of-range value looked like a successful split: for 10 images, a ratio of -0.2 returned 8/2 via negative slicing, and 1.2 (or an accidental percentage like 80) returned 10/0 with no held-out data, no error either way.
train, test = dataset.split(split_ratio=80) # meant 80%
# now raises ValueError naming the problem, instead of silently returning
# every image as train and none as test0 and 1 keep their existing meanings. Not an API break: no signature changed, only previously-silent input now raises.
π Migration guide
No breaking changes in this release.
No deprecations or removals landed in 0.30.6 either. All scheduled remove_in markers in the codebase target 0.31.0 or 0.32.0 and are untouched by this patch release.
Two entries change output values rather than only fixing a crash or a silently-wrong count; not API breaks, but worth checking if your code depends on the old values:
sv.KeyPoints.from_ultralytics/.from_inference/.from_detectron2:as_detections().confidencenow carries the model's own detection score, not the mean of the per-keypoint confidences.sv.DetectionsSmoother:class_idanddatafields on a smoothed track now follow the current frame instead of lagging behind a class change.
π Notable changes
π§ Fixed
Annotators / key points
sv.VertexEllipseHaloAnnotatornow draws the whole halo of a key point whose covariance ellipse is not horizontal. A vertical or diagonal ellipse used to be clipped to a thin band because the fade box was sized as if the major axis always ran along the image x axis. Horizontal ellipses are drawn exactly as before. (#2634)sv.KeyPoints.from_ultralytics,.from_inferenceand.from_detectron2now keep each object's detection score asdetection_confidenceinstead of dropping it, which previously madesv.KeyPoints.with_nmsraiseValueErrorandas_detections()report the mean keypoint confidence instead of the model's own score. (#2633)sv.LabelAnnotatornow sizes the label background correctly for a label containing a blank line. Text drew each blank line at full height while the background measured it as zero, pushing the last line outside the box. Labels without blank lines are unchanged. (#2632)sv.IconAnnotatornow draws a palette icon that carries its own alpha channel (PillowPAmode, storable in TIFF) instead of raisingValueError. Palettes without alpha, and every read with OpenCV installed, are unchanged. (#2624)
Datasets
sv.DetectionDataset.as_yolo/.as_pascal_vocnow write a detection whose mask is empty or has no valid contour as its bounding box instead of silently leaving it out of the label file. (#2631)sv.DetectionDataset.from_yolonow loads a label row carrying a trailing confidence or tracker id (written by Ultralyticssave_txtwithsave_conf=Trueor tracking on) instead of raisingValueError, for both box and segmentation rows. The extra token is ignored; box and polygon geometry are unchanged. An odd-length polygon row with no extra field now loses its last value rather than raising, since coordinate parity is the only signal separating the two cases. (#2619, #2626, #2635)
Tracking
sv.DetectionsSmootherno longer raisesValueError: Conflicting metadatawhen a second tracked object, first seen on a different frame, enters the smoothing window. (#2628)sv.LineZone.triggernow ignores detections with a negativetracker_id(the value trackers such asByteTrackTrackerreport for an unconfirmed track) instead of letting them silently inflatein_count/out_count. (#2623)
Detection utils
sv.polygon_to_masknow accepts a list/tuple/array-like of[x, y]vertices (not just a NumPy array) and returns an all-zero mask for an empty or under-3-vertex polygon instead of crashing inside OpenCV/NumPy with an opaque error. A malformed polygon now raisesValueErrornaming the problem. (#2622)sv.Detections.from_sam3now keeps SAM 3 PVS contour fragments with fewer than 3 vertices instead of dropping them. Single points and 2-point edges are rasterized directly into the mask and bounding box. (#2625)
Image / IO
sv.pillow_to_cv2now converts every Pillow mode to the 8-bit arraycv2.imreadwould produce, reaching every annotator pluscrop_image,resize_image,letterbox_image,scale_image,tint_image,grayscale_image, andplot_image. (#2614)sv.CSVSinknow writes UTF-8 on every platform, preserving non-English detection labels and custom fields on Windows. (#2615)
Docs
- Dead documentation links fixed across the published notebooks (
quickstart.ipynb,annotate-video-with-detections.ipynb,underestand-visitors-with-yolo-world.ipynb). (#2630)
π± Changed
sv.DetectionDataset.split, `sv.ClassificationDatase...
supervision-0.30.5
0.30.5: Tracking, metrics, and image-drawing correctness fixes
supervision 0.30.5 is a patch release closing 11 correctness and crash bugs across line-crossing tracking, mAR@K scoring, model connectors, key points, and image drawing/loading β no new public API of note, no breaking changes. The most consequential fixes are silent, not crashes: LineZone.trigger miscounted crossings by one per flicker whenever a tracker briefly touched the far side of the line, and MeanAverageRecall scored mAR@K against the wrong predictions when a lower-ranked one fit a target more tightly than one within the top K. The remaining fixes close a hard crash in InferenceSlicer on conflicting slice metadata (plus two related Detections.__eq__ bugs), bring the OpenCV-free fallback backend to parity with OpenCV for rotated videos, CMYK images, transparent/1-bit PNGs, and default JPEG/WebP write quality, fix two drawing bugs that reproduce with OpenCV installed (16-bit draw_image, grayscale IconAnnotator icons), and fix two isolated bugs in KeyPoints.as_detections and plot_images_grid. Every fix ships a regression test.
β¨ Spotlights / highlights
sv.LineZone.trigger no longer counts flicker as a crossing
A crossing was confirmed whenever the oldest entry of a minimum_crossing_threshold + 1 frame history differed from every later entry, which never verified the tracker had actually settled on the side it supposedly came from. With minimum_crossing_threshold=2 the side sequence A,A,A,B,A,A,A counted a crossing into A β the side the object never left β so counts drifted by one per flicker, in the wrong direction; three separated flickers gave in_count=3 instead of 0.
line_zone = sv.LineZone(
start=sv.Point(0, 0), end=sv.Point(0, 100), minimum_crossing_threshold=2
)
# a tracker that flickers to the far side for one frame and back
# no longer registers a phantom crossing; only a sustained crossing countsCrossings are now measured against the last side a tracker was confirmed on, not against the oldest history entry. Sustained crossings and minimum_crossing_threshold=1 (the default) are unchanged.
sv.metrics.MeanAverageRecall now scores mAR@K from each image's own top K
The matcher pairs predictions to targets by highest IoU, not confidence, so a prediction ranked below K could take a target away from one ranked within it. Adding a low-confidence duplicate that fit a target more tightly than the top prediction actually lowered mAR@1 β one target with a top prediction at IoU 0.71 scored mAR@1 0.5 alone but 0.0 once a second, unrelated prediction at IoU 1.0 and confidence 0.1 was added. Each detection limit now matches only its own top K predictions.
sv.InferenceSlicer no longer crashes when slices disagree on metadata
RF-DETR and inference-package connectors attach a source_image array per slice; any mismatch across slices previously crashed the merge outright. Any metadata key that isn't identical across every slice is now dropped with a one-time warning naming the dropped keys; source_image is reattached afterward as the full input image.
slicer = sv.InferenceSlicer(callback=callback)
detections = slicer(image) # metadata.source_image is the full input image againFixing this also closed two Detections.__eq__ bugs: NaN in a float-array metadata value now compares equal to itself, and a list-valued value against an ndarray-valued one now returns False instead of raising ValueError.
Rotated videos read upright without OpenCV
OpenCV's FFmpeg backend applies a video's display-rotation matrix to decoded frames and to the reported width/height; the OpenCV-free fallback ignored it, so a portrait phone video came back sideways with its width/height swapped.
sv.KeyPoints.as_detections drops invisible key points from the derived box
as_detections stretched each box over every key point that wasn't [0, 0] or non-finite, ignoring visible β unlike with_nms, which already respected it. A pose whose low-confidence joints were predicted off-frame no longer produces an inflated box.
π Migration guide
No breaking changes in this release.
No deprecations or removals landed in 0.30.5 either. All scheduled remove_in markers in the codebase target 0.31.0 and are untouched by this patch release.
sv.config.SOURCE_IMAGE_METADATA_FIELD is a new public constant (added alongside the InferenceSlicer fix) naming the Detections.metadata key RF-DETR/inference-package connectors use for the source image β additive, no existing code needs to change.
π Notable changes
π§ Fixed
Tracking / metrics
sv.LineZone.triggerno longer counts a spurious crossing in the opposite direction when a tracker flickers to the far side of the line for fewer thanminimum_crossing_thresholdframes. Crossings are now measured against the last side a tracker was confirmed on, and once a tracker has a confirmed side, a new one replaces it only after being held for the full threshold. Sustained crossings,minimum_crossing_threshold=1(the default), and per-tracker isolation are unchanged. (#2600)sv.metrics.MeanAverageRecallnow scores mAR@K from each image's K most confident predictions alone, instead of matching every prediction first and only then keeping the top K by confidence. mAR@1 and mAR@10 now equal the recall of the top 1 and top 10 predictions per image, as documented; mAR@100 changes only for images with more than 100 predictions. (#2604)
Model connectors
sv.InferenceSlicerno longer raises when slices disagree onmetadata(e.g. asource_imageNumPy array attached per-slice by RF-DETR/inference-package connectors), which previously crashed the merge outright. Any metadata key that isn't identical across every slice is now dropped from the merged result, with aSupervisionWarningswarning naming the dropped keys, emitted once per slicer instance.source_imageis a special case: removed from each slice before the lenient merge, then reattached afterward as the full input image. Fixing this also closed two relatedDetections.__eq__bugs: metadata holdingNaNin a float array now compares equal to itself instead of always reading unequal, and comparing a list-valued metadata value against an ndarray-valued one now returnsFalseinstead of raisingValueError. (#2596)
OpenCV-free fallback backend
sv.VideoInfo.from_video_path,sv.get_video_frames_generatorandsv.process_videonow turn a rotated video upright when OpenCV is not installed, as they already do with OpenCV. The fallback now applies quarter and half turns, the same angles OpenCV applies, using the container's display-rotation matrix. Videos without a display rotation, and every read with OpenCV installed, are unchanged. (#2601)sv.IconAnnotatorandsv.draw_imagenow draw CMYK JPEG and TIFF images in their colors when OpenCV is not installed, instead of returning the four ink channels as if they were blue/green/red/alpha. Other images, and every read with OpenCV installed, are unchanged. (#2602)sv.ImageSinkand the dataset exports that encode in-memory images now write JPEG and WebP files at OpenCV's default quality (JPEG 95, lossless WebP) when OpenCV is not installed, instead of Pillow's defaults (JPEG 75, lossy WebP). The fallback's in-memory encoder also now accepts.jpe,.tif,.jp2, and.pgm, which it previously rejected even though file writes already accepted them. PNG and every write with OpenCV installed are unchanged. (#2592)sv.IconAnnotatorandsv.draw_imagenow keep the transparency of grayscale PNGs with alpha, RGB PNGs with a transparent color, and 1-bit PNGs, when OpenCV is not installed β the fallback previously returned Pillow's pixel layout forIMREAD_UNCHANGEDrather than OpenCV's, causing shape-mismatchValueErrors or black-drawn transparent pixels. Other images, and every read with OpenCV installed, are unchanged. (#2588)
Annotators / drawing
sv.draw_imagenow scales a 16-bit PNG down to 8 bits on load, instead of blending it into an 8-bit scene and clipping channel values above 255 β this reproduces with OpenCV installed, sinceIMREAD_UNCHANGEDkeeps a 16-bit PNG at 16 bits. Eight-bit images, and images passed as arrays, are drawn as before. (#2603)sv.IconAnnotatornow draws grayscale PNG icons instead of failing withIndexErrorβ the overlay only handled BGR and BGRA arrays, andcv2.imread(..., IMREAD_UNCHANGED)returns a 2-D array for a grayscale PNG without alpha, with or without OpenCV installed. A grayscale icon is now expanded to BGR on load, and a 16-bit icon is scaled to 8 bits, ascv2.imreaddoes by default. Eight-bit color icons are drawn as before. (#2591)
Key points / utils
sv.KeyPoints.as_detectionsnow leaves key points marked not visible out of the box it derives for each skeleton, assv.KeyPoints.with_nmsalready does. A skeleton with no visible key point is now dropped, as one with only missing key points already was. Key points withoutvisibleconvert as before. (#2605)sv.plot_images_gridnow plots a single image in agrid_size=(1, 1)grid, instead of failing withAttributeErrorβplt.subplotsreturns a loneAxesrather than an array for a 1x1 grid. Grids with more than one cell plot...
supervision-0.30.4
0.30.4: Dataset, connector, and key point correctness fixes
supervision 0.30.4 is a patch release closing 14 correctness and crash bugs across dataset loaders/exporters, model connectors, and key point/annotator/video handling β no new public API, no breaking changes. The most consequential fixes are silent, not crashes: COCO polygon masks loaded shifted by up to a pixel, EXIF-rotated photos loaded with swapped width/height across two loaders and both image backends, and non-ASCII class names were mangled or rejected on Windows across four loaders and one writer. The remaining fixes close hard crashes in the Transformers v4/v5 connectors, Ultralytics pose loading, TraceAnnotator, iterative_seek, and three more dataset-format edge cases. Every fix ships a regression test.
β¨ Spotlights / highlights
sv.DetectionDataset.from_coco no longer shifts mask polygons by up to a pixel
COCO polygons commonly hold sub-pixel float coordinates, but the loader cast them straight to int32, which truncates rather than rounds. Every polygon mask loaded shifted up and to the left by up to a pixel β the same polygon loaded one pixel apart depending on the format it was stored in (a square with corners at 2.6/7.6 covered pixels 2β7 from COCO but 3β8 from LabelMe, an IoU of 0.53 between the two masks). Silent, not a crash β training data was quietly misaligned.
dataset = sv.DetectionDataset.from_coco(
images_directory_path="images",
annotations_path="annotations.json",
) # polygon vertices now rounded to nearest pixel, matching from_yolo/from_labelme/from_pascal_vocA non-finite vertex now raises ValueError naming the annotation id, instead of silently propagating.
from_yolo/as_coco no longer swap width and height on EXIF-rotated photos
Photos from phones are often stored sideways with an EXIF orientation tag. cv2.imread applies that tag, but the size read used Pillow's file-header read, which doesn't. For a quarter-turned photo, from_yolo scaled normalized boxes and polygons by the swapped width/height, so they landed outside the image, and as_coco wrote the swapped dimensions. The OpenCV-free fallback backend now also applies the orientation tag, matching how OpenCV itself handles every read except IMREAD_UNCHANGED β the same file previously loaded with a different shape depending on whether opencv-python was installed.
Dataset loaders/writers no longer mangle non-ASCII class names on Windows
from_coco, from_labelme, from_createml, from_yolo read JSON/YAML, and as_pascal_voc wrote XML, using the platform's default encoding β cp1252 on Windows. A class name like cafΓ© loaded as cafΓΒ©; κ³ μμ΄ raised UnicodeDecodeError. All five now read/write UTF-8 explicitly.
sv.Detections.from_transformers now loads Mask2Former/MaskFormer overlap-safe binary maps
post_process_instance_segmentation(return_binary_maps=True) β the option Transformers recommends when instances can overlap β returns a (num_instances, H, W) stack of binary maps, but the v5 instance path compared it against each segment's id as if it were an id-map, producing a 4-D array mask_to_xyxy rejected outright. Each segment now indexes the stack at its own id, so overlapping instances keep their full masks.
detections = sv.Detections.from_transformers(
transformers_results=processor.post_process_instance_segmentation(
result, target_sizes=[image.size[::-1]], return_binary_maps=True
)[0],
id2label=model.config.id2label,
) # overlapping-instance results now load instead of crashingsv.DetectionDataset.from_pascal_voc no longer silently drops bmp/tif/webp images
The loader only listed .jpg, .jpeg, .png, so every other image was left out without a warning β even though from_yolo/from_folder_structure load those formats and as_pascal_voc writes annotations for them. A dataset exported to Pascal VOC and read back came back smaller than it went out.
π Migration guide
No breaking changes in this release.
No deprecations or removals landed in 0.30.4 either. Eight scheduled-removal commits (ByteTrack, supervision.keypoint, create_tiles/overlay_image, keypoint validators, the legacy MeanAveragePrecision, LMM/from_lmm, and others) plus a new MetricResult ABC exist on develop but are not part of this cherry-picked patch release β they target 0.31.0 and will get their own migration guide when that release ships.
Two internal (non-underscored but unexported) helper functions changed signature as part of fixes in this release: parse_polygon_points (dataset/formats/pascal_voc.py) now returns float64 instead of int, and detections_from_xml_obj's xyxy is now list[list[float]] instead of list[list[int]]. Neither is re-exported from supervision.__init__ or supervision.dataset.__init__, so this is not a public API change β no action needed unless you import these paths directly.
π Notable changes
π§ Fixed
Dataset loaders / exporters
sv.DetectionDataset.from_coconow rounds polygon vertices to the nearest pixel before rasterising masks, instead of truncating them toint32, asfrom_yolo/from_labelme/from_pascal_vocalready do. Truncation shifted every mask up and to the left by up to a pixel β a polygon with corners at2.6/7.6covered pixels 2 to 7 from COCO but 3 to 8 from LabelMe, an IoU of 0.53 between the two masks. A non-finite vertex now raisesValueErrornaming the annotation id; integer vertices and RLE masks load as before. (#2587)sv.DetectionDataset.from_coco/from_labelme/from_createml/from_yolonow read JSON/YAML as UTF-8, andas_pascal_vocwrites XML as UTF-8, instead of the platform default (cp1252 on Windows). A COCO category or YOLOdata.yamlname outside ASCII broke on Windows:cafΓ©loaded ascafΓΒ©,κ³ μμ΄failed withUnicodeDecodeError, andas_pascal_voceither failed withUnicodeEncodeErroror wrote a filefrom_pascal_vocrejected withParseError: not well-formed (invalid token)β Linux/macOS, already UTF-8 by default, are unaffected. (#2585)sv.ClassificationDataset.as_folder_structurenow copies source image files unchanged instead of round-tripping them throughcv2.imread/cv2.imwrite, which dropped alpha channels, downcast 16-bit PNGs to 8-bit, and recompressed JPEGs on every export β matchingsv.DetectionDataset's exports (as_yolo/as_pascal_voc/as_coco), which already copied; exporting into the folder the dataset was loaded from previously rewrote its own images the same lossy way. Images held in memory are still encoded withcv2.imwrite. (#2584)sv.DetectionDataset.from_labelmenow finds images for LabelMe files saved on Windows when loading on Linux or macOS.imagePathis written with\, but the loader took onlyPath(...).name, which doesn't split on\on POSIX β the whole value became the filename and reading failed withValueError: Could not read image from path. The filename is now taken with either separator on every system, as LabelMe itself does when reading its own files; forward-slash paths, and every file on Windows, are unchanged. (#2581)sv.DetectionDataset.from_yolonow loads label files whose class ids are written as decimals (e.g.1.0 0.5 0.5 0.2 0.4), which previously aborted the whole load withValueError: invalid literal for int()βnp.savetxtwrites floats by default and Ultralytics tolerates them. Whole numbers load in any notation now; fractional, non-finite, or non-numeric ids still raise, naming the offending id. (#2580)sv.DetectionDataset.from_yolo/as_coconow size EXIF-oriented images the waycv2.imreadloads them (swapping width/height for orientations 5 to 8), instead of reading the un-rotated file-header size via Pillow β quarter-turned photos previously scaled boxes/polygons by the swapped dimensions and produced mismatched mask shapes. The OpenCV-free fallback backend'simread/imdecodenow apply EXIF orientation too, matching OpenCV's behavior for every read exceptIMREAD_UNCHANGEDβ previously the same file loaded with a different shape depending on whetheropencv-pythonwas installed. (#2577)sv.DetectionDataset.from_pascal_vocno longer fails on annotations with decimal coordinates (e.g.<xmin>48.5</xmin>), which aborted the whole load withValueError: invalid literal for int()β Datumaro, which CVAT uses for its exports, writes VOC this way. Box coordinates are now read as floats and keep their precision; polygon vertices are rounded after the 1-index offset, as the YOLO and LabelMe loaders already do; non-finite values are still rejected. (#2568)sv.DetectionDataset.from_pascal_vocno longer skips.bmp,.tif,.tiff, and.webpimages without a warning β the loader only listed.jpg/.jpeg/.png, even thoughfrom_yolo/from_folder_structureload those formats andas_pascal_vocwrites annotations for them, so a dataset exported to Pascal VOC and read back came back smaller than it went out. It now accepts the same extensions assv.ClassificationDataset.from_folder_structure. (#2569)
Model connectors
sv.Detections.from_transformersnow loads Transformers v5return_binary_maps=Trueinstance results (a(num_instances, H, W)stack), which the v5 path previously compared against each segment'sidas if it were an id-map, producing a 4-D array thatmask_to_xyxyrejected with `ValueError: too many v...
supervision-0.30.3
0.30.3: Pose, VLM, and video/CSV crash and correctness fixes
supervision 0.30.3 is a bug-fix release closing crash and silent-correctness gaps across pose estimation, VLM parsing, video/CSV output, and geometry utilities. Non-finite key points β how pose estimators report an undetected joint β no longer produce duplicate poses that survive sv.KeyPoints.with_nms, or crash the key point annotators outright. sv.Detections.from_vlm now orders backwards box corners, closing a bug where such a box scored a false 0.0 IoU and both survived NMS as a duplicate and counted as a total miss in mAP. sv.TraceAnnotator and sv.CSVSink no longer crash or silently drop columns on the first frame with no detections β a case every non-ByteTrack tracker pipeline hits. sv.process_video no longer hangs forever when max_frames exceeds the video length. Continuing 0.30.2's numeric-correctness theme, sv.pad_boxes and sv.scale_boxes are fixed against integer overflow. No breaking API changes, no new public API.
β¨ Spotlights / highlights
Non-finite key points no longer produce duplicate poses or crash annotators
sv.KeyPoints.with_nms tested key point validity with xy == 0 alone, and NaN β how pose estimators report an undetected joint β is not 0. The stale joint stayed in the NMS box, so a duplicate skeleton scored False on every IoU comparison against it and survived suppression. The key point annotators (sv.VertexAnnotator, sv.EdgeAnnotator, sv.VertexLabelAnnotator, the sv.VertexEllipse*Annotator family) had the matching crash: a single undetected joint raised ValueError: cannot convert float NaN to integer for the whole frame. Both now skip non-finite coordinates, matching sv.KeyPoints.as_detections.
keypoints = sv.KeyPoints(xy=xy, confidence=confidence)
keypoints.with_nms(
threshold=0.5
) # duplicate skeletons with a NaN joint are now suppressedsv.Detections.from_vlm no longer scores a false IoU miss on backwards box corners
A VLM that emits a corner pair backwards produced an xyxy row with x_min > x_max. Nothing downstream caught it: sv.box_iou_batch clamps intersection width at zero, so the box scored 0.0 IoU against itself β surviving NMS as a duplicate and counting as a total miss in mAP β while box_area still reported a plausible positive value. Every VLM parser now orders each box's corners before returning it.
sv.TraceAnnotator and sv.CSVSink no longer crash or silently corrupt output on an empty-detections frame
sv.TraceAnnotator.annotate raised ValueError: The tracker_id field is missing on the first frame with no detections, for every tracker except sv.ByteTrack. Such a frame now draws nothing and still advances the frame counter, so trace_length stays a window over elapsed frames rather than only over populated ones. sv.CSVSink had a quieter failure: an empty batch fixed the CSV header without the data/custom_data columns, and every later row was silently truncated to that schema β dropping fields like class_name for the whole file. The header is now fixed by the first batch that actually carries detections.
sv.process_video no longer hangs forever when max_frames exceeds the video length
The reader thread failed on the out-of-range end before enqueuing its sentinel, leaving the main loop blocked on the read queue indefinitely. max_frames is now capped at the video length, and any reader-thread error surfaces as RuntimeError("Reader thread raised: ...") instead of stalling the call.
Integer-coordinate overflow fixed in sv.pad_boxes and sv.scale_boxes
Both computed intermediate values that could overflow or silently wrap for large integer coordinates (e.g. large int32/uint16/int64 boxes). Both now use overflow-safe arithmetic.
xyxy = np.array([[10, 20, 30, 40]], dtype=np.int64)
sv.pad_boxes(xyxy=xyxy, px=5, py=10) # int64 output, no wraparoundpad_boxes changes return dtype for integer input β see the migration guide below.
sv.scale_image and sv.resize_image(keep_aspect_ratio=True) no longer crash on an extreme aspect ratio or tiny scale factor
A small enough factor β or an aspect ratio too extreme for the target box β could round an output axis down to 0, and cv2.resize raised an assertion naming nothing the caller passed. Each axis now keeps at least one pixel. Two callers inherit the fix: sv.letterbox_image could not fill the resolution it was asked for, and sv.CropAnnotator with scale_factor < 1 aborted the whole frame as soon as one detection box was a few pixels across.
π Migration guide
No breaking API changes. One fix changes return dtype for integer input:
sv.pad_boxesβ integerxyxynow returnsint64(orfloat64if a padded coordinate exceeds theint64range), instead of the input's original integer dtype, which could silently overflow or wrap for small dtypes likeint16/uint8.
If your code assumes pad_boxes preserves the input's exact dtype (e.g. reusing the result as an int16 array), cast explicitly: sv.pad_boxes(...).astype(np.int16).
sv.scale_boxes also fixes an integer-overflow bug, but its return dtype was already float64 for integer input before this release β unaffected.
π Notable changes
π§ Fixed
sv.Detections.from_ultralyticsnow assigns the placeholder class ID0to every mask in a masks-only result, instead of sequential IDs across masks that belong to the same image. (#2566)sv.pad_boxesnow computes integer-coordinate padding without overflow or unsigned casting errors. (#2565)sv.scale_imageandsv.resize_image(keep_aspect_ratio=True)no longer derive a zero-sized target, fixing a crash reached viasv.letterbox_imageandsv.CropAnnotator. (#2564)sv.KeyPoints.with_nmsno longer stops suppressing duplicate skeletons as soon as a key point is non-finite. (#2563)sv.tint_imageno longer tints the caller's own image array in place. (#2562)sv.LineZoneno longer consumes thetriggering_anchorsiterable during validation, so a generator ormappassed in is no longer exhausted before the firsttrigger()call. (#2561)- The key point annotators now skip key points whose coordinates are not finite instead of raising
ValueError. (#2560) sv.PolygonZonenow rejects a polygon with fewer than three vertices instead of building a zone that can never trigger;sv.Detections.from_vlmnow orders each parsed box's corners. (#2554)sv.ClassificationDataset.as_folder_structurenow rejects images that would overwrite the same class-relative filename before writing any files. (#2551)sv.process_videono longer hangs forever whenmax_framesis larger than the number of frames in the video. (#2546)sv.filter_polygons_by_areaandsv.approximate_polygonnow preserve local geometry for large-origin integer andfloat64polygons. (#2542)sv.TraceAnnotator.annotateno longer raises on an empty-detections frame;sv.CSVSinkno longer lets an empty batch fix the CSV header. (#2539)sv.scale_boxesnow preserves exact integer intermediates, preventing overflow and scaled-corner rounding errors for large integer-coordinate boxes. (#2541)- A release's own version-pinned docs no longer show the outdated-version banner on the day it ships, and the docs-publish and canonical-backfill workflows now share one
gh-pageswrite lock instead of racing each other. (#2536)
π Contributors
- kevin (@kevin9327) β fixed
KeyPoints.with_nms/key point annotators crashing on non-finite key points,tint_imageimage aliasing,LineZonegenerator exhaustion, and thescale_image/resize_imagezero-target crash - Durgamani Sasikumar (@tedo001, LinkedIn) β fixed
PolygonZone/from_vlmbox-corner ordering and theTraceAnnotator/CSVSinkempty-frame crash - S B Pranay (@pranaysb, LinkedIn) β fixed integer overflow in
scale_boxesand dtype loss infilter_polygons_by_area/approximate_polygon - JiantaoPeng (@PengJianT) β fixed
from_ultralyticsmasks-only class ID sizing - trueoneplusone (@trueoneplusone) β fixed integer overflow in
pad_boxes - Andrew Barnes (@Bortlesboat, LinkedIn) β fixed classification export filename collisions
- Abhijith Neil Abraham (@abhijithneilabraham, LinkedIn) β fixed
process_videohanging whenmax_framesexceeds the video length - Jirka Borovec (@Borda, LinkedIn) β fixed the release-day outdated-docs banner
Full changelog: 0.30.2...0.30.3
supervision-0.30.2
0.30.2: Detection numeric-correctness fixes
supervision 0.30.2 fixes three silent numeric-correctness bugs in the detection utilities β integer box areas that could wrap negative on large boxes, and two coordinate converters that truncated fractional values on integer input β plus an InferenceSlicer determinism fix that restores its documented source-order result guarantee under multithreading. A set of versioned-docs reliability fixes rounds out the release. No breaking API changes, no new public API.
β¨ Spotlights / highlights
sv.Detections.box_area no longer overflows to a negative number
Integer-coordinate box area now computes in float64. A large int32 box (50000 x 50000) previously wrapped to a negative area.
detections = sv.Detections(xyxy=np.array([[0, 0, 50000, 50000]], dtype=np.int32))
detections.box_area # array([2.5e+09]) β was negative before the fixxcycwh_to_xyxy / denormalize_boxes stop truncating integer boxes
Both converters wrote fractional half-extent or scaled coordinates into a copy of the integer input, silently truncating toward zero β relevant when converting quantized VLM output (e.g. boxes on a 0..1000 grid).
xcycwh_to_xyxy(np.array([[10, 10, 5, 5]], dtype=np.int32))
# array([[ 7.5, 7.5, 12.5, 12.5]]) β the fractional coordinate 7.5 is no longer truncated to 7InferenceSlicer merges results in source order under multithreading
Slice results now merge in source order under thread_workers > 1, restoring the ordering guarantee its docstring documents. Row order β and, for tied confidences, which overlapping box survives with_nms/with_nmm β no longer varies between runs on identical input.
π Migration guide
No breaking API changes, but three fixes above change return dtype for integer input:
sv.Detections.box_area/.areaβ integerxyxynow returnsfloat64(was the input's integer dtype, which could silently overflow)sv.xcycwh_to_xyxyβ integer input now returnsfloat64(was truncated integer output)sv.denormalize_boxesβ integer input now returnsfloat64(was truncated integer output)
If your code indexes arrays with these outputs (e.g. image[y1:y2, x1:x2]), a float64 result raises TypeError: slice indices must be integers. Cast explicitly where integer indices are required: xcycwh_to_xyxy(boxes).astype(int).
π Notable changes
π§ Fixed
sv.Detections.box_area(andsv.Detections.areafor axis-aligned boxes) now computes integer-coordinate box areas infloat64, preventing integer overflow for large boxes. (#2514)sv.xcycwh_to_xyxyno longer truncates coordinates for integer input arrays. (#2515)sv.denormalize_boxesno longer truncates coordinates for integer input arrays. (#2516)sv.InferenceSlicernow merges slice results in source order whenthread_workers > 1, restoring the ordering guarantee its docstring documents. (#2517)- Docs deployment for
latestno longer fails witherror: version 'latest' already existswhenlatestexists as an alias of a released version. (#2512, #2513) - Versioned documentation builds now emit a valid
/latest/search/SearchAction URL when Mike removes the trailing slash fromsite_url; docs CI renders the custom theme under Mike version contexts to protect the URL, version banners, and star JSON-LD. Applies to future builds going forward. (#2529) - Versioned documentation banners now adjust MkDocs Material's desktop sidebar inline layout and scroll height without shifting the mobile navigation drawer. (#2532)
- Versioned documentation deploys now export the version being built, so the outdated-version banner reaches readers of the
developtree; the workflow also now backs up the pre-rewritegh-pagestip to a timestamped branch before committing over it. Applies to future builds going forward. (#2533) - The canonical-backfill workflow now reports (in the job summary) rewritten canonicals whose target page does not exist under
latest/, and backfills the outdated-version banner itself into already-published archive trees, patching the empty banner markup those pages already carry rather than rebuilding them. (#2534)
π Contributors
- Advait Shukla (@AdvaitS) β fixed integer overflow in
box_area - Guillaume Flambard (@guillaume-flambard, LinkedIn) β fixed integer truncation in
denormalize_boxesandxcycwh_to_xyxy - Roshan Sharma (@roshaninfordham, LinkedIn) β fixed
InferenceSlicerresult ordering under multithreading - Jirka Borovec (@Borda, LinkedIn) β versioned-docs banner layout,
MIKE_DOCS_VERSIONexport,gh-pagesbackup, and canonical-backfill fixes (#2529, #2532, #2533, #2534)
Full changelog: 0.30.1...0.30.2
supervision-0.30.1
0.30.1: Numeric-precision and stability fixes
supervision 0.30.1 is a bug-fix patch release. It corrects numeric-precision issues that only surface on specific inputs β large-coordinate oriented boxes (geospatial data, stitched frames), large integer boxes for box_iou, and rotated tracks in DetectionsSmoother β where prior versions could silently return imprecise or self-inconsistent results instead of erroring. It also fixes a duplicate-libavdevice-load crash risk on macOS when both av and opencv-python are installed, plus smaller fixes to list_files_with_extensions and the cv2-free RGBA fallback. No public API was added or removed, and no signature changed β a drop-in upgrade from 0.30.0 for virtually all users. See Migration guide below for the one narrow exception (box_iou on complex-valued coordinates) and for the precision caveats on the numeric fixes.
β¨ Spotlights / highlights
1. Oriented-box area/IoU precision fix for large coordinates
sv.Detections.area and sv.oriented_box_iou_batch now translate OBB coordinates to a local origin before floating-point math. Previously, large-coordinate inputs could lose enough precision that a box's IoU with itself collapsed below 1.0.
pair_origin = np.minimum(origin_i, origin_j)
offset_i = (origin_i - pair_origin).astype(np.float32, copy=False)
offset_j = (origin_j - pair_origin).astype(np.float32, copy=False)2. sv.box_iou no longer overflows on large integer boxes
Area computation now takes coordinate differences before casting to float, avoiding int32 overflow. For realistic coordinate magnitudes, box_iou's scalar result now matches box_iou_batch.
3. Duplicate libavdevice crash fixed on macOS
import supervision no longer loads PyAV's native libraries when the OpenCV backend is active β PyAV is now imported lazily, only where it's used, preventing a duplicate libavdevice warning (and possible crash) when both av and opencv-python are installed.
4. DetectionsSmoother keeps oriented-box corners consistent
Smoothed OBB corners are now aligned (start index + winding) to a reference before averaging, so rotated tracks smooth correctly instead of averaging mismatched corner orderings.
5. sv.get_polygon_center precision fix for large-coordinate polygons
Centroid calculation now translates to the first vertex and computes in float64 before adding the origin back, preventing integer overflow and precision loss for realistic coordinate magnitudes.
π Migration guide
No public signature changed. One item below (box_iou on complex coordinates) does make one specific previously-succeeding call now raise β narrow and deliberate, not classified as breaking since complex-valued box coordinates were never a documented/supported input. The rest only change output values for inputs that were already edge cases:
sv.box_iouon complex-valued coordinates: previously silently discarded the imaginary part and returned a real number. Now raisesTypeError("box coordinates must be real-valued").- OBB precision fixes: results for oriented boxes with large coordinates or rotated tracks may differ slightly from 0.30.0 β the new values are the corrected ones. Re-calibrate any hardcoded IoU/area thresholds tuned against the old (imprecise) output.
sv.box_iou/box_iou_batchagreement: for realistic integer coordinate magnitudes (below 2^53),box_iou's scalar result now matchesbox_iou_batch. Not a universal guarantee βbox_iousubtracts before casting to float,box_iou_batchstill casts tofloat64before subtracting, so the two can diverge at coordinates β₯ 2^53 (~9 quadrillion), far outside any real use case.
π Notable changes
π Added
- RF-DETR example scripts (
rfdetr_example.py) added to thecount_people_in_zone,heatmap_and_track,speed_estimation,tracking, andtraffic_analysisbundled examples. (#2497)
π± Changed
sv.box_iounow raisesTypeErrorfor complex-valued box coordinates instead of silently discarding the imaginary part. (#2485)- Performance:
DetectionsSmoother.update_with_detectionsnow checks active tracker IDs via set membership instead of scanning per tracked object. No output changes. (#2496) - Documentation and API-reference examples now default to RF-DETR instead of Ultralytics YOLO. (#2493, #2494, #2497)
π§ Fixed
- RF-DETR speed estimation now measures elapsed source-frame intervals, including gaps when tracked detections are temporarily missed. (#2497)
sv.get_polygon_centernow calculates polygon centroids in translatedfloat64coordinates, preventing integer overflow and precision loss for realistic-magnitude large-coordinate polygons. (#2491)sv.Detections.areaandsv.oriented_box_iou_batchnow translate oriented-box coordinates to local origins before floating-point math, preventing self-IoU collapse for large-coordinate inputs. (#2492)DetectionsSmoothernow keeps oriented-box corners aligned with smoothedxyxygeometry, including rotated tracks and mixed metadata windows. (#2489)sv.box_iounow calculates overlap infloat64, preventingint32area overflow for large boxes; its scalar result now matchessv.box_iou_batchfor realistic coordinate magnitudes. (#2485)sv.list_files_with_extensionsno longer includes directories when listing all files without an extension filter. (#2486)sv.pillow_to_cv2now accepts RGBA images when the cv2-free fallback backend is active, matching OpenCV by dropping alpha and returning BGR channels. (#2488)import supervisionno longer loads PyAV's native libraries when the OpenCV backend is selected; PyAV is now imported lazily on first use, preventing a duplicatelibavdevicewarning (and possible crash) on macOS when bothavandopencv-pythonare installed. (#2509)- Docstring examples converted to executed doctests across
annotators/core.py,dataset/formats/coco.py,dataset/formats/createml.py, andDetections.from_vlm; previously wrong documented outputs corrected for several VLM examples. (#2474, #2475, #2479, #2484)
Also in this release: routine dependency bumps (dependabot: astral-sh/setup-uv, wheel, pymdown-extensions x2, pypa/gh-action-pypi-publish, twine, cryptography), CI/docs-workflow maintenance, and test-only additions (geometry contract test, sklearn parity test) β none change installed package behavior. (#2028, #2470, #2472, #2473, #2480, #2481, #2482, #2483, #2499, #2501, #2506, #2507, #2508)
π Contributors
- lawliet (@lawliet206) β oriented-box area/IoU precision fix at large coordinate origins
- Zhewen Tan (@tandede) β polygon centroid overflow/precision fix
- Tamil Adhavan S K (@adhavan18, LinkedIn) β DetectionsSmoother oriented-box corner alignment fix; added geometry contract test
- NIKHIL (@Nikhi00718) β fixed
box_iouint32 overflow; added complex-coordinateTypeErrorguard - shao (@shaoming11, LinkedIn) β
DetectionsSmoothertracker-ID lookup performance improvement - Tyyyy (@uczltw6) β RGBA image support in the cv2-free fallback conversion
- ZZZZZ (@BruceWae) β fixed
list_files_with_extensionsto exclude directories - FootysHands (@ayo0la) β converted docstring examples to doctests; corrected wrong documented VLM example outputs
- Swapnil Gautam (@Swapnil-gautam) β converted docstring examples to doctests in dataset format modules
- Daniiiil1 (@Daniiiil1) β added sklearn parity test for metrics
- Christoph Deil (@cdeil, LinkedIn) β repo maintenance docs
Full changelog: 0.30.0...0.30.1
supervision-0.30.0
v0.30.0: Run supervision without OpenCV
supervision 0.30.0 makes OpenCV optional. A new private _cv2/ backend (NumPy and Pillow, with PyAV for the video path) reimplements every OpenCV call the library needs, so supervision now runs on opencv-python-headless β or no OpenCV wheel at all β instead of crashing on import. This release also adds Soft-NMS, LabelMe and CreateML dataset formats, GeoTIFF-aware windowed reads for InferenceSlicer, and ships five breaking changes, most notably OpenCV no longer being installed by default, JSONSink switching to native JSON types, and mask_non_max_merge computing exact mask overlap instead of a downscaled approximation. Python 3.9 support is dropped β 3.10 is now the minimum.
β¨ Spotlights / highlights
Run supervision without OpenCV
import supervision as sv
window = sv.ImageWindow("frame")
for frame in sv.get_video_frames_generator("input.mp4"):
window.show(frame)
if window.wait_key(1) == "q":
breakThe largest change in this release: OpenCV stays the default backend when installed, but supervision no longer requires it β there's no opencv-python extra anymore either. sv.ImageWindow replaces cv2.imshow/cv2.waitKey for display. av>=14.2 is now a required dependency for the PyAV video path during this transition. See the OpenCV migration guide.
Soft-NMS
detections = sv.Detections.from_ultralytics(result)
softened = detections.with_soft_nms(sigma=0.5)
filtered = detections.with_soft_nms(sigma=0.5, score_threshold=0.3)sv.Detections.with_soft_nms (plus sv.box_soft_non_max_suppression / sv.mask_soft_non_max_suppression) rescales overlapping detections' confidence instead of discarding them outright β useful in crowded scenes where hard NMS drops valid overlapping objects.
New dataset formats + GeoTIFF-aware, batched slicing
dataset = sv.DetectionDataset.from_labelme(
images_directory_path="images/",
annotations_directory_path="annotations/",
)
import rasterio
with rasterio.open("RGB.byte.tif") as raster:
slicer = sv.InferenceSlicer(callback=my_model_callback, batch_size=4)
detections = slicer(raster)DetectionDataset.from_labelme/as_labelme and from_createml/as_createml join the existing COCO/YOLO/Pascal-VOC converters. sv.InferenceSlicer can now read an open rasterio dataset window-by-window for multi-GB aerial/drone GeoTIFFs without loading the whole image (pip install "supervision[geotiff]"), and accepts batch_size for batched-callback inference.
sv.load_image_from_url
image = sv.load_image_from_url("https://media.roboflow.com/notebooks/examples/dog.jpeg")Load an image straight from an HTTP(S) URL as an OpenCV array, with optional on-disk caching.
π Migration guide
Five breaking changes. Most require no code changes beyond a type check or threshold recalibration. The two that need action from most users: the OpenCV install change below, and the Python 3.10 floor.
OpenCV is no longer installed by default. If a compatible cv2 is already importable in your environment, nothing changes for you β it's still preferred automatically. Otherwise install one wheel family yourself (pip install opencv-python or opencv-python-headless) if you need OpenCV-specific behavior, then restart the process β cv2 is detected once at import time. sv.ImageWindow replaces cv2.imshow/cv2.waitKey. Full guide: docs/how_to/opencv_migration.md.
Python 3.10+ is now required β 3.9 reached end-of-life in October 2025.
sv.JSONSink now emits native JSON types, not strings:
# before 0.30.0
row["score"] == "0.85" # str
row["is_valid"] == "True" # str
# after 0.30.0
row["score"] == 0.85 # float
row["is_valid"] is True # boolsv.CSVSink stays textual, but its per-row custom-data slicing now matches JSONSink.
sv.mask_non_max_merge computes exact mask overlap, not a downscaled approximation, and ignores the now-deprecated mask_dimension parameter (kept for signature compatibility, removal in 0.33.0). Re-tune your overlap threshold after upgrading. Passing overlap_metric/mask_dimension positionally still works β the values are still honored β but now emits a DeprecationWarning; pass them by keyword to silence it. More than five positional arguments raises TypeError.
Detections.merge() on mixed dense + CompactMask inputs now returns a CompactMask, not a plain ndarray:
merged = sv.Detections.merge([dense_detections, compact_mask_detections])
isinstance(merged.mask, np.ndarray) # was True, now False β it's a CompactMaskOnly affects code that explicitly merges a CompactMask-carrying Detections object with a dense-mask one yourself β InferenceSlicer, DetectionsSmoother, and with_nms/with_nmm always merge type-homogeneous lists internally, so they're unaffected. The all-dense merge path is also unchanged. This is a substantial performance win: ~2500Γ less peak memory, ~13Γ faster on a 1080p frame with 40 detections. If you need the old return type without touching every call site: call merged.mask = merged.mask.to_dense() right after merge(), or avoid producing CompactMask in the first place (Detections.from_inference(compact_masks=False), the default).
supervision also now requires av>=14.2 as an install-time dependency for the PyAV cv2-free video path β this doesn't change any API, so it isn't counted as breaking, but pinned/vendored environments should account for it.
Deprecation removals pushed back one release: ByteTrack, supervision.keypoint, normalized_xyxy, and supervision.dataset.utils RLE compatibility shims β originally scheduled for removal in 0.30.0 β are now scheduled for 0.31.0 instead, giving a full transition window.
π Notable changes
π Added
sv.load_image_from_urlβ load an HTTP(S) image as an OpenCV array, with optional on-disk caching (#2372)- cv2-free PyAV video fallback + private
_cv2backend facade β image/geometry/drawing/text/video without OpenCV (#2430, #2431, #2432, #2433, #2435, #2438, #2439, #2440, #2441, #2443) sv.ImageWindowβ tkinter+Pillow desktop window replacingcv2.imshow/cv2.waitKey(#2320)- Soft-NMS β
sv.box_soft_non_max_suppression,sv.mask_soft_non_max_suppression,sv.Detections.with_soft_nms(#1624) sv.VLM.GOOGLE_GEMINI_3_5βDetections.from_vlmparses Gemini 3.5 output (#2449)get_video_frames_generator(prefetch=...)β background-thread decode into a bounded queue (#2273)PolygonZone(require_all_anchors=...)β toggle all-anchors vs. any-anchor containment (#2272)KeyPoints.merge()β combine a list ofKeyPoints, mirroringDetections.merge(#2412)BaseAnnotator.requires_maskβ class-level flag on all annotators (#2370)CompactMask.from_coco_rle+Detections.from_inference(compact_masks=True)(#2367)CompactMask.image_shapeproperty (#2383)sv.mask_to_roiβ exclusive mask-bound helper for slicing/crops (#2416)DetectionDataset.from_labelme/as_labelme(#2299)DetectionDataset.from_createml/as_createml(#2284)InferenceSlicerGeoTIFF support βsv.WindowedRasterDataset,pip install "supervision[geotiff]"(#2281)InferenceSlicer(batch_size=...)β batched callback contract (#1239)ConfusionMatrix.benchmark(save_directory_path=...)β adaptive TP/FP/FN validation-mosaic export (#2271)HeatMapAnnotator.reset(),TraceAnnotator.reset(),DetectionsSmoother.reset()β clear accumulated per-stream state, so a single instance can be reused across independent streams (#2418)AREA_DATA_FIELDconfig constant (#2428)sv.denormalize_boxesandsv.xyxyxyxy_to_xyxynow exported at the top level
β οΈ Breaking Changes
- OpenCV no longer installed by default; no OpenCV extra (#2443)
- Python 3.10+ required β 3.9 dropped (#2260, #2381)
sv.JSONSinkemits native JSON types instead of strings;sv.CSVSinkcustom-data slicing now matchesJSONSink(#2400)sv.mask_non_max_mergecomputes exact overlap, ignoresmask_dimension, positionaloverlap_metric/mask_dimensiondeprecated (#2400)Detections.merge()on mixed dense +CompactMaskinputs returnsCompactMask(#2383)
π± Changed
DetectionDataset/ClassificationDatasetequality now compares orderedclasseslists, not an unordered setsupervisionnow requiresav>=14.2as an install-time dependency for the cv2-free video fallback β no API change (#2438)- Deprecation-window delays:
ByteTrack,supervision.keypoint,normalized_xyxy, dataset-utils RLE compat removals moved0.30.0β0.31.0 - Perf:
count_nonzeromask pixel counts (#2361), vectorizedbox_iou_batch_with_jaccard(#2359), faster mask-annotation ROI blending (#2368), fewer corner circles on square label backgrounds (#2346), less compact-mask materialization in the polygon annotator (#2369) - Geometry-aware IoU/area dispatch centralized (#2374)
π§ Fixed
sv.Recalltracks prediction-only classes, matchingPrecision/F1Score(#2467, #2468)DetectionDataset.from_pascal_vocno longer raises on background images, with or withoutforce_masks=True(#2463, #2469)import supervisionno longer surfaces the deprecatedByteTrackwarning- Reopening
sv.CSVSink/sv.JSONSinkstarts a fresh session β no stale rows or header (#2459) from_vlmGemini 2.0/2.5/3.5 salvages valid entries from partially malformed JSON arrays (#2449)save_coco_annotations/as_cocoread image sizes from headers, no pixel decode for labels-only export (#2442)sv.F1Scoreno longer emits a spurious div-by-zero `RuntimeWarni...
supervision-0.29.1
What's new
π KeyPoints.with_nms() β NMS for pose estimation
import supervision as sv
key_points = model.predict(image) # sv.KeyPoints
key_points = key_points.with_nms(threshold=0.5) # removes duplicate skeletonsDerives axis-aligned bounding boxes from each skeleton's valid (non-zero and visible) keypoints, then applies standard box NMS. Supports class_agnostic mode and any OverlapMetric (IOU, IOS). Raises ValueError if detection_confidence is not set.
feat-keypoints-with-nms-compressed.mp4
Notable changes
Bug fixes
-
sv.DetectionDataset.as_pascal_vocno longer mutates bounding boxes (#2341) Previously, every export shifted every bounding box by +1 px in-place. A second call compounded the shift. Fixed by rebinding to a new array; on-disk XML output is unchanged. -
sv.Precisionandsv.F1Scorecorrectly count background false positives (#2331) Predictions on images with no ground-truth objects, and predictions of classes absent from any annotation, were previously ignored. UnderMICROandMACROaveraging they are now counted as false positives.WEIGHTEDaveraging is unchanged. Users should re-evaluate existing metric results after upgrading. -
sv.DetectionsSmootherworks with confidence-free detections (#2333) The smoother no longer raises when detections have no confidence scores. Confidence is averaged over the frames that carry it; tracks without any confidence produceNone. -
sv.Detections.from_vlmis robust to malformed Gemini/Qwen output (#2342) Valid JSON that is not a list, or whose elements are not dicts, now degrades to emptyDetectionsinstead of raisingTypeError. A malformed mask value in Gemini 2.5 responses no longer misaligns thexyxy/confidence/masksarrays. -
sv.JSONSinkserializes NumPy scalars incustom_data(#2334)np.int64frame indices and other NumPy scalars incustom_datano longer raiseTypeErrorat flush time. NumPy arrays are serialized as lists. The file handle closes even when serialization fails. -
sv.approximate_polygonrespects the point-count budget (#2332) The function now returns at mostfloor(N * (1 - percentage))points (minimum 3). Previously it could return more points than requested.epsilon_stepis now validated to be positive. -
COCO export preserves all segments for multi-part masks (#2322) Previously, only the first polygon was written when a non-crowd detection had disjoint mask segments. All polygon parts are now written.
Performance
-
sv.HaloAnnotatoris ~4Γ faster withCompactMaskdetections (#2339)HaloAnnotatornow uses the same optimized CompactMask paint path asMaskAnnotator. Previously it materialized each mask full-frame; now it operates on the bounding-box crop. Annotated output is unchanged. -
Mask IoU uses less peak memory (#2323) Mask IoU computation now uses matrix multiplication on flattened masks instead of an explicit
(N, M, H, W)tensor. For masks larger than 4096Γ4096 px, computation promotes to float64 automatically. Results are numerically identical. -
sv.mask_to_xyxyandsv.KeyPoints.as_detectionsvectorized (#2330) Both functions now use batched NumPy operations instead of per-element loops. Outputs are bit-identical.
Contributors
- Ruben Haisma (@RubenHaisma, LinkedIn) β VLM robustness, Pascal VOC export fix, DetectionsSmoother, JSONSink, metrics correctness, polygon budgeting, vectorization
- Agis Kounelis (@kounelisagis, LinkedIn) β HaloAnnotator perf, mask IoU matmul, mask_to_xyxy/KeyPoints.as_detections vectorization, OBB cookbook
- Piotr Skalski (@SkalskiP, LinkedIn) β
KeyPoints.with_nms() - Abdelrahman Gomaa (@abdogomaa201099, LinkedIn) β COCO multi-polygon export
Full Changelog: 0.29.0...0.29.1