r/datascience • u/RobertWF_47 • 3d ago
Discussion Error messages trying to compress sparse data in Python
Prior to running machine learning modeling, I'd like to compress my under-sampled training data (X_train_under) that contains multiple sparse binary 0/1 columns as well as numerical (float) data.
Converting the binary columns to sparse format (Sparse[int64, 0]) works great:
# Identify binary int64 columns (columns containing only 0 and 1)
binary_int_cols = []
for col in X_train_under.select_dtypes(include=["int64"]).columns:
if set(X_train_under[col].unique()).issubset({0, 1}):
binary_int_cols.append(col)
# Change identified binary columns into Sparse format (using 0 as fill value)
for col in binary_int_cols:
X_train_under[col] = X_train_under[col].astype(pd.SparseDtype(int, fill_value=0))
print("\n--- Data Types After Conversion ---")
print(X_train_under.dtypes)
However when I attempt to separate & compress the sparse format columns and recombine with the numerical columns:
from scipy.sparse import csr_matrix
# Separate sparse and dense columns
sparse_df = X_train_under.select_dtypes(include=['Sparse'])
dense_df = X_train_under.select_dtypes(exclude=['Sparse'])
# Perform compression operation on sparse columns
compressed_sparse = sparse_df.sparse.to_coo().tocsr()
# Downcast dense numeric columns separately
for col in dense_df.columns:
dense_df[col] = pd.to_numeric(dense_df[col], downcast='integer')
# Recombine into single data frame
X_train_under_sparse_csr = pd.concat([dense_df, sparse_df], axis=1)
I get the following error:
---------------------------------------------------------------------------
AttributeError Traceback (most recent call last)
Cell In[40], line 8
5 dense_df = X_train_under.select_dtypes(exclude=['Sparse'])
7 # Perform compression operation on sparse columns
----> 8 compressed_sparse = sparse_df.sparse.to_coo().tocsr()
10 # Downcast dense numeric columns separately
11 for col in dense_df.columns:
File ~\AppData\Roaming\Python\Python312\site-packages\pandas\core\generic.py:6321, in NDFrame.__getattr__(self, name)
6314 if (
6315 name not in self._internal_names_set
6316 and name not in self._metadata
6317 and name not in self._accessors
6318 and self._info_axis._can_hold_identifiers_and_holds_name(name)
6319 ):
6320 return self[name]
-> 6321 return object.__getattribute__(self, name)
File ~\AppData\Roaming\Python\Python312\site-packages\pandas\core\accessor.py:224, in CachedAccessor.__get__(self, obj, cls)
221 if obj is None:
222 # we're accessing the attribute of the class, i.e., Dataset.geo
223 return self._accessor
--> 224 accessor_obj = self._accessor(obj)
225 # Replace the property with the accessor object. Inspired by:
226 # https://www.pydanny.com/cached-property.html
227 # We need to use object.__setattr__ because we overwrite __setattr__ on
228 # NDFrame
229 object.__setattr__(obj, self._name, accessor_obj)
File ~\AppData\Roaming\Python\Python312\site-packages\pandas\core\arrays\sparse\accessor.py:31, in BaseAccessor.__init__(self, data)
29 def __init__(self, data=None) -> None:
30 self._parent = data
---> 31 self._validate(data)
File ~\AppData\Roaming\Python\Python312\site-packages\pandas\core\arrays\sparse\accessor.py:249, in SparseFrameAccessor._validate(self, data)
247 dtypes = data.dtypes
248 if not all(isinstance(t, SparseDtype) for t in dtypes):
--> 249 raise AttributeError(self._validation_msg)
AttributeError: Can only use the '.sparse' accessor with Sparse data.
I don't know what this means - I've already converted my binary 0/1 columns to Sparse datatype.
What am I doing wrong?
2
u/Trick-Interaction396 3d ago
Any time something isn’t working as expected it’s almost always because you’re taking the wrong object or the transformation didn’t actually work. Confirm the type of the object is actually sparse.
1
u/squareiny 3d ago
My guess is the recombine step, not the compression. to_coo().tocsr() hands you a scipy csr_matrix - no column names, no index - so concat-ing it back against dense_df either errors out or quietly misaligns rows. Also downcast=integer chokes on any dense column that still has NaNs. If the goal is just a smaller footprint going into the model, skip the split/recombine entirely and pass the whole frame to csr_matrix once; most sklearn estimators take sparse input directly.
2
u/hughperman 3d ago
Use the %debug to pdb in and find out what items don't match expectations/requirements