Currently, projects are extracted by aligning them to an empty Original versification corpus. This allows extracts to be (mostly) reliably aligned line-by-line across project versifications. While addressing some other problems, I realized that the Original versification does not actually include all possible canonical books including some of those supported in the other base versifications (e.g. LAO which exists in the English). This means that projects with content in these books will never see that content used as training data when using silnlp (Serval doesn't have this problem). I don't expect this happens frequently or perhaps at all, but it is something we should keep in mind and perhaps look to solve as we have time.
Currently, projects are extracted by aligning them to an empty Original versification corpus. This allows extracts to be (mostly) reliably aligned line-by-line across project versifications. While addressing some other problems, I realized that the Original versification does not actually include all possible canonical books including some of those supported in the other base versifications (e.g. LAO which exists in the English). This means that projects with content in these books will never see that content used as training data when using silnlp (Serval doesn't have this problem). I don't expect this happens frequently or perhaps at all, but it is something we should keep in mind and perhaps look to solve as we have time.