tidymodels/butcherPublic

NotificationsYou must be signed in to change notification settings
Fork16
Star137

Reduce the size of model objects saved to disk

License

Unknown, MIT licenses found

Licenses found

137 stars 16 forks Branches Tags Activity

Star

Notifications

You must be signed in to change notification settings

Branches Tags

Folders and files

Name		Name	Last commit message	Last commit date
Latest commit History 792 Commits
.github		.github
.vscode		.vscode
R		R
inst		inst
man		man
revdep		revdep
tests		tests
vignettes		vignettes
.Rbuildignore		.Rbuildignore
.covrignore		.covrignore
.gitignore		.gitignore
DESCRIPTION		DESCRIPTION
LICENSE		LICENSE
LICENSE.md		LICENSE.md
NAMESPACE		NAMESPACE
NEWS.md		NEWS.md
README.Rmd		README.Rmd
README.md		README.md
_pkgdown.yml		_pkgdown.yml
air.toml		air.toml
butcher.Rproj		butcher.Rproj
codecov.yml		codecov.yml
cran-comments.md		cran-comments.md

Repository files navigation

butcher

Overview

Modeling or machine learning in R can result in fitted model objectsthat take up too much memory. There are two main culprits:

Heavy usage of formulas and closures that capture the enclosingenvironment in model training
Lack of selectivity in the construction of the model object itself

As a result, fitted model objects contain components that are oftenredundant and not required for post-fit estimation activities. Thebutcher package provides tooling to “axe” parts of the fitted outputthat are no longer needed, without sacrificing prediction functionalityfrom the original model object.

Installation

Install the released version from CRAN:

install.packages("butcher")

Or install the development version fromGitHub:

# install.packages("pak")pak::pak("tidymodels/butcher")

Butchering

As an example, let’s wrap anlm model so it contains a lot ofunnecessary stuff:

library(butcher)our_model<-function() {some_junk_in_the_environment<- runif(1e6)# we didn't know about  lm(mpg~.,data=mtcars)}

This object is unnecessarily large:

library(lobstr)obj_size(our_model())#> 8.02 MB

When, in fact, it should only be:

small_lm<- lm(mpg~.,data=mtcars)obj_size(small_lm)#> 22.22 kB

To understand which part of our original model object is taking up themost memory, we leverage theweigh() function:

big_lm<- our_model()weigh(big_lm)#> # A tibble: 25 × 2#>    object            size#>    <chr>            <dbl>#>  1 terms         8.01#>  2 qr.qr         0.00666#>  3 residuals     0.00286#>  4 fitted.values 0.00286#>  5 effects       0.0014#>  6 coefficients  0.00109#>  7 call          0.000728#>  8 model.mpg     0.000304#>  9 model.cyl     0.000304#> 10 model.disp    0.000304#> # ℹ 15 more rows

The problem here is in theterms component of ourbig_lm. Because ofhowlm() is implemented in thestats package, the environment inwhich our model was made is carried along in the fitted output. Toremove the (mostly) extraneous component, we can usebutcher():

cleaned_lm<- butcher(big_lm,verbose=TRUE)#> ✔ Memory released: 8.00 MB#> ✖ Disabled: `print()`, `summary()`, and `fitted()`

Comparing it against oursmall_lm, we find:

weigh(cleaned_lm)#> # A tibble: 25 × 2#>    object           size#>    <chr>           <dbl>#>  1 terms        0.00771#>  2 qr.qr        0.00666#>  3 residuals    0.00286#>  4 effects      0.0014#>  5 coefficients 0.00109#>  6 model.mpg    0.000304#>  7 model.cyl    0.000304#>  8 model.disp   0.000304#>  9 model.hp     0.000304#> 10 model.drat   0.000304#> # ℹ 15 more rows

And now it will take up about the same memory on disk assmall_lm:

weigh(small_lm)#> # A tibble: 25 × 2#>    object            size#>    <chr>            <dbl>#>  1 terms         0.00763#>  2 qr.qr         0.00666#>  3 residuals     0.00286#>  4 fitted.values 0.00286#>  5 effects       0.0014#>  6 coefficients  0.00109#>  7 call          0.000728#>  8 model.mpg     0.000304#>  9 model.cyl     0.000304#> 10 model.disp    0.000304#> # ℹ 15 more rows

To make the most of your memory available, this package provides five S3generics for you to remove parts of a model object:

axe_call(): To remove the call object.
axe_ctrl(): To remove controls associated with training.
axe_data(): To remove the original training data.
axe_env(): To remove environments.
axe_fitted(): To remove fitted values.

When you runbutcher(), you execute all of these axing functions atonce. Any kind of axing on the object will append a butchered class tothe current model object class(es) as well as a new attribute namedbutcher_disabled that lists any post-fit estimation functions that aredisabled as a result.

Model Object Coverage

Check out thevignette("available-axe-methods") to see butcher’scurrent coverage. If you are working with a new model object that couldbenefit from any kind of axing, we would love for you to make a pullrequest! You can visit thevignette("adding-models-to-butcher") formore guidelines, but in short, to contribute a set of axe methods:

Runnew_model_butcher(model_class = "your_object", package_name = "your_package")
Use butcher helper functionsweigh() andlocate() to decide whatto axe
Finalize edits toR/your_object.R andtests/testthat/test-your_object.R
Make a pull request!

Contributing

This project is released with aContributor Code ofConduct.By contributing to this project, you agree to abide by its terms.

For questions and discussions about tidymodels packages, modeling, andmachine learning, pleasepost on RStudioCommunity.
If you think you have encountered a bug, pleasesubmit anissue.
Either way, learn how to create and share areprex(a minimal, reproducible example), to clearly communicate about yourcode.
Check out further details oncontributing guidelines for tidymodelspackages andhow to gethelp.

About

Reduce the size of model objects saved to disk

butcher.tidymodels.org/

Resources

Readme

License

Unknown, MIT licenses found

Releases13

butcher 0.4.0 Latest

Dec 9, 2025

+ 12 releases

Packages

No packages published

Contributors18

+ 4 contributors

Movatterモバイル変換

Navigation Menu

Search code, repositories, users, issues, pull requests...

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

License

Licenses found

Uh oh!

Folders and files

Latest commit

History

Repository files navigation

butcher

Overview

Installation

Butchering

Model Object Coverage

Contributing

About

Resources

License

Licenses found

Code of conduct

Uh oh!

Stars

Watchers

Forks

Releases13

Packages

Uh oh!

Contributors18

Uh oh!

Languages

Movatterモバイル変換

License

Licenses found

tidymodels/butcher

Folders and files

Latest commit

History

Repository files navigation

butcher

Overview

Installation

Butchering

Model Object Coverage

Contributing

About

Resources

License

Licenses found

Code of conduct

Uh oh!

Stars

Watchers

Forks

Releases13

Packages0

Uh oh!

Contributors18

Uh oh!

Languages

Packages