Skip to content

Question about reading public listings on soujava/vagas-java (personal final-year project) #1956

Description

@manuchehrqoriev798

Hello soujava/vagas-java team,

I'm Manuchehr Qoriev, a computer science student. OmniLadders (https://omniladders.pages.dev) is my own personal final-year project, free and non-commercial, and not run by any university or organization. It helps students around the world find internships and graduate roles.

My script already reads your public job listings, normally once a day, and I would like to check that this is OK with you. For each job the site shows the title, employer, location and a link to the job, and the date where your list gives one, plus a few things derived from the advert (tags, a visa note, a fit score). The same list is also a public file anyone can download. Nothing is sold. This repository has no licence file, so I am asking rather than assuming, and I would be glad to credit it with a link back.

Could you tell me if this is fine, or what you would prefer (a lower rate, an API or feed, only some fields)? If someone else handles this, please forward it.

If you would rather I stop, tell me. I will block it within a few days, so that from the next daily run after that there are no more requests and your listings leave the site, and I will delete the advert text I stored.

Thank you,
Manuchehr Qoriev
OmniLadders · omniladders@gmail.com · https://omniladders.pages.dev

Technical details, in case someone asks:

  • Requests, once a day (1 in all; a failed request is retried at most twice):
    GET api.github.com/repos/soujava/vagas-java/issues?state=open&per_page=100 (the open issues, one request)
    They send the headers of Chrome on macOS; I can switch to "Mozilla/5.0 (compatible; OmniLadders/1.0; +https://omniladders.pages.dev)" if you prefer.
  • Link check: once a day, a HEAD request to the link of each job shown on the site (your job page or, where a listing provides one, the employer's own application page), following redirects (plus a GET on 403, 404, 405, 410 or a server error), as "Mozilla/5.0 (compatible; OmniLadders/1.0; +https://omniladders.pages.dev)".
  • Stored privately: the advert text, where your site provides it, used to work out labels (visa, degree, experience, language); not published in full. An occasional manual check may send, at most, a job's title, employer, location, link and up to about 2,000 characters of its text to a third-party AI service (Google Gemini or OpenAI).
  • Published: title, employer, location, date, salary if given, a few tags derived from the advert, a visa note, a computed fit score and a link (also as https://omniladders.pages.dev/jobs.json), plus a short quote where an advert states a visa, experience or language requirement.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions