Issue #8 – Add Dataset Preparation Workflow with Git LFS and Extraction Script - #13
Conversation
Co-authored-by: Sarah Gershuni <sarah556726@gmail.com> Co-authored-by: Ruti Cohen <r0583283575@gmail.com>
Co-authored-by: Sarah Gershuni <sarah556726@gmail.com> Co-authored-by: Ruti Cohen <r0583283575@gmail.com>
Co-authored-by: Sarah Gershuni <sarah556726@gmail.com> Co-authored-by: Ruti Cohen <r0583283575@gmail.com>
Co-authored-by: Sarah Gershuni <sarah556726@gmail.com> Co-authored-by: Ruti Cohen <r0583283575@gmail.com>
|
Please add a quick reference in |
|
Code Review: Download and Prepare Dataset Great job on automating the dataset preparation workflow! Strengths:
Suggestions for improvement:
Overall: |
| @@ -0,0 +1,15 @@ | |||
| # Dataset Preparation (Git LFS) | |||
There was a problem hiding this comment.
Currently, there is no scripts/prepare_data.sh file, which is required by the DoD and mentioned in the usage instructions.
This script should:
- Validate that Git LFS is installed.
- Run
git lfs pull. - Invoke the Python script.
This ensures the end-to-end setup works seamlessly after a fresh clone.
| git clone <repo-url> myproj && cd myproj | ||
| git lfs pull # fetch real bytes | ||
| ./scripts/prepare_data.sh # extract every archive in dataset/ | ||
| ``` No newline at end of file |
There was a problem hiding this comment.
The documentation could be more complete.
Please consider adding:
- Dataset details (size ~10.67 GB, 8,034 images, structure overview).
- An example using the new
prepare_data.shscript, since it's easier for users than running the Python script directly.
| from pathlib import Path | ||
|
|
||
| DATASET_DIR = Path("dataset") # Source of archives + extraction destination | ||
| OVERWRITE_DIR = True # Automatic deletion if directory already exists |
There was a problem hiding this comment.
The script currently always deletes existing directories (OVERWRITE_DIR=True).
To make it safer and more flexible, please add an option like --force for overwriting.
If .prepared exists and --force is not set, the script should skip extraction instead of deleting data.
There was a problem hiding this comment.
I would not continue now with the python implementation, it can be done more easily using bash
| This code fetches the real dataset files stored with Git LFS and then extracts all archives in the dataset/ folder, so the project is ready to use right after cloning. | ||
|
|
||
| ```bash | ||
| git clone <repo-url> myproj && cd myproj |
There was a problem hiding this comment.
Let's not mention this command because
- we already have this repository
- it's missleading because if I copy the command I will create repository "myproj" which is not needed
There was a problem hiding this comment.
Please rename this file to include full words, e.g. dataset_preparation.md
| @@ -0,0 +1,2 @@ | |||
| *.tar.gz filter=lfs diff=lfs merge=lfs -text | |||
There was a problem hiding this comment.
are we using *.tag.gz files? If not, let's remove it
There was a problem hiding this comment.
The data preparation script could also be implemented as a bash script, I don't see the need for python here.
I see only 1 archive called "images.zip", I don't know if there's everything within (annotations etc).
The easy bash script (I asked chatgpt to generate it) would be:
#!/usr/bin/env bash
set -o errexit
set -o nounset
set -o pipefail
set -o xtrace
# Base directories
script_dir=$(dirname "$(realpath "$0")")
workspace_dir=${script_dir}
dataset_dir=${workspace_dir}/dataset
# Ensure dataset exists
[[ -d ${dataset_dir} ]] || { echo "dataset/ not found" >&2; exit 1; }
# Expect exactly one .zip
archive=(${dataset_dir}/*.zip)
[[ ${#archive[@]} -eq 1 ]] || { echo "Expected exactly one zip archive in dataset/"; exit 1; }
arc=${archive[0]}
base=$(basename "$arc")
out_dir=${dataset_dir}/${base%.zip}
# Skip if already extracted
if [[ -d ${out_dir} ]]; then
echo "Skipping ${arc} (output dir exists)"
exit 0
fi
mkdir -p "${out_dir}"
echo "Extracting ${base} → ${out_dir}"
unzip -q "${arc}" -d "${out_dir}"
echo "ok" > "${out_dir}/.prepared"
echo "Finished ${base}"
| from pathlib import Path | ||
|
|
||
| DATASET_DIR = Path("dataset") # Source of archives + extraction destination | ||
| OVERWRITE_DIR = True # Automatic deletion if directory already exists |
There was a problem hiding this comment.
I would not continue now with the python implementation, it can be done more easily using bash
There was a problem hiding this comment.
are there only images in this archive?
Please install the tree utility (https://tree.readthedocs.io/en/latest/) and share here the dir structure (w/o files in dirs) after extraction. Also add this sctructure to the README
There was a problem hiding this comment.
Right now, we have uploaded only a few images in one archive, just to ensure that our script works properly.
Is it necessary to upload a folder with the same structure as the dataset?
…iles Co-authored-by: Sarah Gershuni <sarah556726@gmail.com> Co-authored-by: Ruti Cohen <r0583283575@gmail.com>
Co-authored-by: Sarah Gershuni <sarah556726@gmail.com> Co-authored-by: Ruti Cohen <r0583283575@gmail.com>
Co-authored-by: Sarah Gershuni <sarah556726@gmail.com> Co-authored-by: Ruti Cohen <r0583283575@gmail.com>
Co-authored-by: Sarah Gershuni <sarah556726@gmail.com> Co-authored-by: Ruti Cohen <r0583283575@gmail.com>
…et_prep.py and map_dataset.py for improved readability.
Issue #8 – Add Dataset Preparation Workflow with Git LFS and Extraction Script
Description
This PR introduces the dataset preparation workflow required for handling large archives via Git LFS.
The workflow ensures datasets can be pulled and extracted in a fully automated way after a fresh clone.
Changes
dataset/directory for archives (.zip,.tar,.tar.gz,.tgz,.tar.bz2,.tar.xz).git lfs pull)..preparedmarker file on success..zipand.tar.gzfiles to be tracked via Git LFS.Usage
After cloning the repository and pulling LFS data: