Skip to content

Issue #8 – Add Dataset Preparation Workflow with Git LFS and Extraction Script - #13

Merged
Sarah5567 merged 22 commits into
mainfrom
dataset-preparation
Aug 28, 2025
Merged

Sarah5567 merged 22 commits into
mainfrom
dataset-preparation

Conversation

@Sarah5567

Copy link
Copy Markdown
Collaborator

Description

This PR introduces the dataset preparation workflow required for handling large archives via Git LFS.
The workflow ensures datasets can be pulled and extracted in a fully automated way after a fresh clone.

Changes

  • Added scripts/prepare_data.py:
    • Scans the dataset/ directory for archives (.zip, .tar, .tar.gz, .tgz, .tar.bz2, .tar.xz).
    • Validates that files are not Git LFS pointers (requires git lfs pull).
    • Safely extracts archives into named subdirectories.
    • Creates a .prepared marker file on success.
    • Supports overwrite of existing directories.
  • Added docs/dataset_prep.md:
    • Short usage instructions.
    • Fresh-clone workflow example.
  • Configured .gitattributes for .zip and .tar.gz files to be tracked via Git LFS.
  • Added logging and path-traversal protection for secure extraction.

Usage

After cloning the repository and pulling LFS data:

git lfs pull
./scripts/prepare_data.sh

Sarah5567 and others added 4 commits August 25, 2025 21:07
Co-authored-by: Sarah Gershuni <sarah556726@gmail.com>
Co-authored-by: Ruti Cohen <r0583283575@gmail.com>
Co-authored-by: Sarah Gershuni <sarah556726@gmail.com>
Co-authored-by: Ruti Cohen <r0583283575@gmail.com>
Co-authored-by: Sarah Gershuni <sarah556726@gmail.com>
Co-authored-by: Ruti Cohen <r0583283575@gmail.com>
Co-authored-by: Sarah Gershuni <sarah556726@gmail.com>
Co-authored-by: Ruti Cohen <r0583283575@gmail.com>
@Sarah5567 Sarah5567 self-assigned this Aug 25, 2025
@Rut-Vahab

Copy link
Copy Markdown
Collaborator

Please add a quick reference in README.md pointing to docs/dataset_prep.md so that users know how to set up the dataset after cloning.

@chani0343

Copy link
Copy Markdown
Collaborator

Code Review: Download and Prepare Dataset

Great job on automating the dataset preparation workflow!

Strengths:

  • Using Git LFS for large archives is best practice and keeps the repo manageable.
  • The extraction script (prepare_data.sh/py) is clear and ensures data is pulled before extraction.
  • Documentation in docs/dataset_prep.md is concise and explains the workflow.
  • The dataset structure is well described and the script maintains directory organization.

Suggestions for improvement:

  • Consider adding error handling in the script for missing archives or failed extraction.
  • Add a check to verify that Git LFS is installed and initialized before running extraction.
  • If possible, print a summary (number of images, folders extracted) at the end of the script for user feedback.
  • In the docs, clarify the expected directory structure after extraction (maybe with a tree diagram).
  • If the dataset is very large, suggest in the docs how to download/extract only a subset for quick tests.

Overall:
Well done! The workflow is reproducible and easy to follow. Just a few minor improvements for robustness and user experience.

@chani0343 chani0343 closed this Aug 26, 2025
@chani0343 chani0343 reopened this Aug 26, 2025
@chani0343 chani0343 closed this Aug 26, 2025
@chani0343 chani0343 reopened this Aug 26, 2025
Comment thread docs/dataset_prep.md Outdated
@@ -0,0 +1,15 @@
# Dataset Preparation (Git LFS)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Currently, there is no scripts/prepare_data.sh file, which is required by the DoD and mentioned in the usage instructions.
This script should:

  • Validate that Git LFS is installed.
  • Run git lfs pull.
  • Invoke the Python script.

This ensures the end-to-end setup works seamlessly after a fresh clone.

Comment thread docs/dataset_prep.md Outdated
git clone <repo-url> myproj && cd myproj
git lfs pull # fetch real bytes
./scripts/prepare_data.sh # extract every archive in dataset/
``` No newline at end of file

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The documentation could be more complete.
Please consider adding:

  • Dataset details (size ~10.67 GB, 8,034 images, structure overview).
  • An example using the new prepare_data.sh script, since it's easier for users than running the Python script directly.

Comment thread scripts/prepare_data.py Outdated
from pathlib import Path

DATASET_DIR = Path("dataset") # Source of archives + extraction destination
OVERWRITE_DIR = True # Automatic deletion if directory already exists

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The script currently always deletes existing directories (OVERWRITE_DIR=True).
To make it safer and more flexible, please add an option like --force for overwriting.
If .prepared exists and --force is not set, the script should skip extraction instead of deleting data.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I would not continue now with the python implementation, it can be done more easily using bash

Repository owner deleted a comment from chani0343 Aug 26, 2025
Comment thread docs/dataset_prep.md Outdated
This code fetches the real dataset files stored with Git LFS and then extracts all archives in the dataset/ folder, so the project is ready to use right after cloning.

```bash
git clone <repo-url> myproj && cd myproj

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Let's not mention this command because

  • we already have this repository
  • it's missleading because if I copy the command I will create repository "myproj" which is not needed

Comment thread docs/dataset_prep.md Outdated

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Please rename this file to include full words, e.g. dataset_preparation.md

Comment thread .gitattributes Outdated
@@ -0,0 +1,2 @@
*.tar.gz filter=lfs diff=lfs merge=lfs -text

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

are we using *.tag.gz files? If not, let's remove it

Comment thread scripts/prepare_data.py Outdated

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The data preparation script could also be implemented as a bash script, I don't see the need for python here.

I see only 1 archive called "images.zip", I don't know if there's everything within (annotations etc).

The easy bash script (I asked chatgpt to generate it) would be:

#!/usr/bin/env bash
set -o errexit
set -o nounset
set -o pipefail
set -o xtrace

# Base directories
script_dir=$(dirname "$(realpath "$0")")
workspace_dir=${script_dir}
dataset_dir=${workspace_dir}/dataset

# Ensure dataset exists
[[ -d ${dataset_dir} ]] || { echo "dataset/ not found" >&2; exit 1; }

# Expect exactly one .zip
archive=(${dataset_dir}/*.zip)
[[ ${#archive[@]} -eq 1 ]] || { echo "Expected exactly one zip archive in dataset/"; exit 1; }

arc=${archive[0]}
base=$(basename "$arc")
out_dir=${dataset_dir}/${base%.zip}

# Skip if already extracted
if [[ -d ${out_dir} ]]; then
    echo "Skipping ${arc} (output dir exists)"
    exit 0
fi

mkdir -p "${out_dir}"
echo "Extracting ${base} → ${out_dir}"

unzip -q "${arc}" -d "${out_dir}"

echo "ok" > "${out_dir}/.prepared"
echo "Finished ${base}"

Comment thread scripts/prepare_data.py Outdated
from pathlib import Path

DATASET_DIR = Path("dataset") # Source of archives + extraction destination
OVERWRITE_DIR = True # Automatic deletion if directory already exists

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I would not continue now with the python implementation, it can be done more easily using bash

Comment thread dataset/images.zip Outdated

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

are there only images in this archive?
Please install the tree utility (https://tree.readthedocs.io/en/latest/) and share here the dir structure (w/o files in dirs) after extraction. Also add this sctructure to the README

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Right now, we have uploaded only a few images in one archive, just to ensure that our script works properly.
Is it necessary to upload a folder with the same structure as the dataset?

Sarah5567 and others added 4 commits August 27, 2025 23:17
Co-authored-by: Sarah Gershuni <sarah556726@gmail.com>
Co-authored-by: Ruti Cohen <r0583283575@gmail.com>
…et_prep.py and map_dataset.py for improved readability.
@Sarah5567
Sarah5567 merged commit 05a9bb4 into main Aug 28, 2025
2 checks passed
@r83575 r83575 self-assigned this Sep 8, 2025
r83575 pushed a commit that referenced this pull request Sep 8, 2025
Issue #8 – Add Dataset Preparation Workflow with Git LFS and Extraction Script
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants