Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
17 changes: 17 additions & 0 deletions .github/workflows/build-test.yml
Original file line number Diff line number Diff line change
Expand Up @@ -46,13 +46,30 @@ jobs:
shell: bash

- uses: prefix-dev/setup-pixi@v0.10.2
if: matrix.environment != 'ci-sklearn-nightly'
with:
pixi-version: v0.68.0
environments: ${{ matrix.environment }}
# we can freeze the environment and manually bump the dependencies to the
# latest version time to time.
frozen: true

# The nightly environment is installed by hand: the locked pandas nightly
# wheel is dropped from its index after a few days, so it is refreshed
# first. The action accepts no other input together with run-install.
- uses: prefix-dev/setup-pixi@v0.10.2
if: matrix.environment == 'ci-sklearn-nightly'
with:
pixi-version: v0.68.0
run-install: false

- name: Install the nightly environment with the latest pandas nightly
if: matrix.environment == 'ci-sklearn-nightly'
run: |
pixi update --environment ci-sklearn-nightly pandas
pixi install --environment ci-sklearn-nightly --frozen
shell: bash

- name: Linters
run: pixi run -e lint lint

Expand Down
16 changes: 16 additions & 0 deletions docs/changes.rst
Original file line number Diff line number Diff line change
Expand Up @@ -17,6 +17,22 @@ v0.17
with an error while it is read, before anything in it is audited or
constructed, instead of failing with an unrelated error, or being accepted,
during construction. :pr:`547` by `Adrin Jalali`_.
- Add support for pandas objects: :class:`~pandas.DataFrame`,
:class:`~pandas.Series`, every kind of :class:`~pandas.Index`, and the
extension arrays and extension dtypes of pandas itself, not those of other
libraries, can now be saved and loaded. They are stored as
the numpy arrays and scalars they are made of and rebuilt through the public
pandas constructors, so no pandas internals end up in the file, and they are
trusted by default, except for the pyarrow backed arrays and dtypes. The
one dtype that cannot be saved is ``ArrowDtype(pyarrow.string())``, whose
name pandas reserves for its ``StringDtype``.
Estimators from other libraries that keep pandas objects
in their fitted attributes, such as ``category_encoders``, can now be
persisted. A file written with one pandas version loads with any other from
2.0 on, keeping the dtypes of the version that wrote it. Not preserved are
the ``freq`` of datetime-like indexes and arrays, the ``attrs`` and ``flags``
of a Series or DataFrame, and the storage, python or pyarrow, of a string
dtype. :pr:`552` by `Adrin Jalali`_.
- Fix a regression since v0.12.0 where saving an object whose ``__reduce__``
raises failed at dump time. ``__reduce__`` is called on every object to
detect a plain constructor call, but Cython extension types with a
Expand Down
15 changes: 14 additions & 1 deletion docs/persistence.rst
Original file line number Diff line number Diff line change
Expand Up @@ -250,7 +250,20 @@ Supported libraries
Skops intends to support all of **scikit-learn**, that is, not only its
estimators, but also other classes like cross validation splitters. Furthermore,
most types from **numpy** and **scipy** should be supported, such as (sparse)
arrays, dtypes, random generators, and ufuncs.
arrays, dtypes, random generators, and ufuncs. **pandas** objects, that is
``DataFrame``, ``Series``, every kind of ``Index``, and the extension arrays
and extension dtypes of pandas itself, not those of other libraries, are
supported as well with pandas 2.0 or later: they are
stored as the arrays they are made of and rebuilt through the public pandas
constructors, so that no pandas internals end up in the file, and a file
written with one pandas version loads with any other. They are trusted by
default, except for the pyarrow backed arrays and dtypes, which need to be
passed as ``trusted`` explicitly; ``ArrowDtype(pyarrow.string())`` cannot be
saved at all, since its name is the one pandas reserves for ``StringDtype``.
Not preserved are the
``freq`` of datetime-like indexes and arrays, the ``attrs`` and ``flags`` of a
``Series`` or ``DataFrame``, and the storage, python or pyarrow, of a string
dtype, which is an environment choice over the same values.

Apart from this core, we plan to support machine learning libraries commonly
used be the community. So far, we have tested the following libraries:
Expand Down
2 changes: 1 addition & 1 deletion docs/requirements.txt
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
# to be synced with the versions in pyproject.toml
matplotlib>=3.3
pandas>=1
pandas>=2
fairlearn>=0.7.0
sphinx>=3.2.0
sphinx-gallery>=0.7.0
Expand Down
1,110 changes: 1,039 additions & 71 deletions pixi.lock

Large diffs are not rendered by default.

32 changes: 27 additions & 5 deletions pyproject.toml
Original file line number Diff line number Diff line change
Expand Up @@ -131,7 +131,6 @@ skops = { path = ".", editable = true }
[tool.pixi.feature.docs.dependencies]
# To be synced with the versions in docs/requirements.txt
matplotlib = ">=3.3"
pandas = ">=1"
sphinx = ">=3.2.0"
sphinx-gallery = ">=0.7.0"
sphinx-rtd-theme = ">=1"
Expand All @@ -141,8 +140,10 @@ sphinx-issues = ">=1.2.0"

[tool.pixi.feature.docs.pypi-dependencies]
# everything that depends on scikit-learn needs to be a pypi dependency so that this
# spec is compatible with the nightly build environment.
# spec is compatible with the nightly build environment. The same holds for pandas,
# whose dev version the nightly build environment installs from pypi.
fairlearn = ">=0.7.0"
pandas = ">=2"

[tool.pixi.feature.tests.dependencies]
pytest = ">=7"
Expand All @@ -151,20 +152,28 @@ flaky = ">=3.7.0"
pandoc = ">=3.6.4"
rich = ">=12"
matplotlib = ">=3.3"
pandas = ">=1"

[tool.pixi.feature.tests.pypi-dependencies]
# these are packages that require scikit-learn. They need to be as a pypi dependency
# because otherwise there will be a package resolution conflict between pypi and conda
# when installing pre-release nightly release.
lightgbm = ">=3"
xgboost = ">=1.6"
# skops.io supports pandas 2.0 and later; each CI environment pins one minor
# version so that the whole range is tested, see the sklearn* features below.
# A pypi dependency for the same reason as above: the nightly environment
# installs the dev version of pandas from pypi.
pandas = ">=2"

[tool.pixi.feature.lint.dependencies]
pre-commit = "*"

[tool.pixi.feature.dev.dependencies]
ipython = "*"
# A conda pandas for the default environment, which includes the nightly index
# of the sklearn-nightly feature: that index shadows PyPI for the packages it
# has, so a pypi pandas there would be a dev build.
pandas = ">=2"

[tool.pixi.feature.sklearn12.dependencies]
scikit-learn = "~=1.2.0"
Expand Down Expand Up @@ -242,6 +251,8 @@ numpy = "~=2.5.0"
scipy = "~=1.18.0"
catboost = ">=1.0"
quantile-forest = "~=1.4.0"
# keeps pandas objects in fitted attributes, see the test for issue #450
category_encoders = ">=2.6"
python = "~=3.14.0"

# [tool.pixi.feature.sklearn17]
Expand All @@ -260,10 +271,21 @@ extra-index-urls = ["https://pypi.anaconda.org/scientific-python-nightly-wheels/
# The version value here needs to be exact, hence == instead of ~=
scikit-learn = "==1.10.dev0"
fairlearn = "*"
pandas = "*"
numpy = "*"
scipy = "*"

[tool.pixi.feature.pandas-nightly.pypi-dependencies]
# The dev version of pandas from the nightly index; "*" would pick the latest
# release, since pre-releases are only considered when named explicitly. Unlike
# the scikit-learn and scipy nightlies, pandas puts the git revision in the
# version, so every build is a new file and the index drops old ones after a
# few days: the locked wheel goes stale. CI refreshes it before installing the
# environment, see .github/workflows/build-test.yml; do the same locally with
# ``pixi update --environment ci-sklearn-nightly pandas``. This is why the
# feature is not part of the default environment. Like the scikit-learn pin
# above, the version needs a bump after each pandas release.
pandas = "==3.2.0.dev0"

[tool.pixi.feature.lint.tasks]
lint = { cmd = "pre-commit install && pre-commit run -v --all-files --show-diff-on-failure" }

Expand All @@ -282,4 +304,4 @@ ci-sklearn16 = ["rich", "tests", "lint", "sklearn16"]
ci-sklearn17 = ["rich", "tests", "lint", "sklearn17"]
ci-sklearn18 = ["rich", "tests", "lint", "sklearn18"]
ci-sklearn19 = ["rich", "tests", "lint", "sklearn19"]
ci-sklearn-nightly = ["rich", "tests", "lint", "sklearn-nightly"]
ci-sklearn-nightly = ["rich", "tests", "lint", "sklearn-nightly", "pandas-nightly"]
Loading
Loading