tidymodels/textrecipesPublic

NotificationsYou must be signed in to change notification settings
Fork17
Star164

Extra recipes for Text Processing

License

Unknown, MIT licenses found

Licenses found

164 stars 17 forks Branches Tags Activity

Star

Notifications

You must be signed in to change notification settings

Branches Tags

Folders and files

Name		Name	Last commit message	Last commit date
Latest commit History 867 Commits
.github		.github
.vscode		.vscode
R		R
data-raw		data-raw
data		data
man-roxygen		man-roxygen
man		man
pkgdown/favicon		pkgdown/favicon
revdep		revdep
src		src
tests		tests
vignettes		vignettes
.Rbuildignore		.Rbuildignore
.clang-format		.clang-format
.gitignore		.gitignore
DESCRIPTION		DESCRIPTION
LICENSE		LICENSE
LICENSE.md		LICENSE.md
NAMESPACE		NAMESPACE
NEWS.md		NEWS.md
README.Rmd		README.Rmd
README.md		README.md
_pkgdown.yml		_pkgdown.yml
air.toml		air.toml
codecov.yml		codecov.yml
cran-comments.md		cran-comments.md
textrecipes.Rproj		textrecipes.Rproj

Repository files navigation

textrecipes

Introduction

textrecipes contain extra steps for therecipes package forpreprocessing text data.

Installation

You can install the released version of textrecipes fromCRAN with:

install.packages("textrecipes")

Install the development version from GitHub with:

# install.packages("pak")pak::pak("tidymodels/textrecipes")

Example

In the following example we will go through the steps needed, to converta character variable to the TF-IDF of its tokenized words after removingstopwords, and, limiting ourself to only the 10 most used words. Thepreprocessing will be conducted on the variablemedium andartist.

library(recipes)library(textrecipes)library(modeldata)data("tate_text")okc_rec<- recipe(~medium+artist,data=tate_text)|>  step_tokenize(medium,artist)|>  step_stopwords(medium,artist)|>  step_tokenfilter(medium,artist,max_tokens=10)|>  step_tfidf(medium,artist)okc_obj<-okc_rec|>  prep()str(bake(okc_obj,tate_text))#> tibble [4,284 × 20] (S3: tbl_df/tbl/data.frame)#>  $ tfidf_medium_colour     : num [1:4284] 2.31 0 0 0 0 ...#>  $ tfidf_medium_etching    : num [1:4284] 0 0.86 0.86 0.86 0 ...#>  $ tfidf_medium_gelatin    : num [1:4284] 0 0 0 0 0 0 0 0 0 0 ...#>  $ tfidf_medium_lithograph : num [1:4284] 0 0 0 0 0 0 0 0 0 0 ...#>  $ tfidf_medium_paint      : num [1:4284] 0 0 0 0 2.35 ...#>  $ tfidf_medium_paper      : num [1:4284] 0 0.422 0.422 0.422 0 ...#>  $ tfidf_medium_photograph : num [1:4284] 0 0 0 0 0 0 0 0 0 0 ...#>  $ tfidf_medium_print      : num [1:4284] 0 0 0 0 0 ...#>  $ tfidf_medium_screenprint: num [1:4284] 0 0 0 0 0 0 0 0 0 0 ...#>  $ tfidf_medium_silver     : num [1:4284] 0 0 0 0 0 0 0 0 0 0 ...#>  $ tfidf_artist_akram      : num [1:4284] 0 0 0 0 0 0 0 0 0 0 ...#>  $ tfidf_artist_beuys      : num [1:4284] 0 0 0 0 0 ...#>  $ tfidf_artist_ferrari    : num [1:4284] 0 0 0 0 0 0 0 0 0 0 ...#>  $ tfidf_artist_john       : num [1:4284] 0 0 0 0 0 0 0 0 0 0 ...#>  $ tfidf_artist_joseph     : num [1:4284] 0 0 0 0 0 ...#>  $ tfidf_artist_león       : num [1:4284] 0 0 0 0 0 0 0 0 0 0 ...#>  $ tfidf_artist_richard    : num [1:4284] 0 0 0 0 0 0 0 0 0 0 ...#>  $ tfidf_artist_schütte    : num [1:4284] 0 0 0 0 0 0 0 0 0 0 ...#>  $ tfidf_artist_thomas     : num [1:4284] 0 0 0 0 0 0 0 0 0 0 ...#>  $ tfidf_artist_zaatari    : num [1:4284] 0 0 0 0 0 0 0 0 0 0 ...

Breaking changes

As of version 0.4.0,step_lda() no longer accepts character variablesand instead takes tokenlist variables.

the following recipe

recipe(~text_var,data=data)|>  step_lda(text_var)

can be replaced with the following recipe to achive the same results

lda_tokenizer<-function(x)text2vec::word_tokenizer(tolower(x))recipe(~text_var,data=data)|>  step_tokenize(text_var,custom_token=lda_tokenizer  )|>  step_lda(text_var)

Contributing

This project is released with aContributor Code ofConduct.By contributing to this project, you agree to abide by its terms.

For questions and discussions about tidymodels packages, modeling, andmachine learning, pleasepost on RStudioCommunity.
If you think you have encountered a bug, pleasesubmit anissue.
Either way, learn how to create and share areprex(a minimal, reproducible example), to clearly communicate about yourcode.
Check out further details oncontributing guidelines for tidymodelspackages andhow to gethelp.

About

Extra recipes for Text Processing

textrecipes.tidymodels.org/

Resources

Readme

License

Unknown, MIT licenses found

Releases20

textrecipes 1.1.0 Latest

Mar 19, 2025

+ 19 releases

Packages

No packages published

Movatterモバイル変換

Navigation Menu

Search code, repositories, users, issues, pull requests...

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

License

Licenses found

Uh oh!

Folders and files

Latest commit

History

Repository files navigation

textrecipes

Introduction

Installation

Example

Breaking changes

Contributing

About

Resources

License

Licenses found

Code of conduct

Uh oh!

Stars

Watchers

Forks

Releases20

Packages

Uh oh!

Contributors12

Uh oh!

Languages

Movatterモバイル変換

License

Licenses found

tidymodels/textrecipes

Folders and files

Latest commit

History

Repository files navigation

textrecipes

Introduction

Installation

Example

Breaking changes

Contributing

About

Resources

License

Licenses found

Code of conduct

Uh oh!

Stars

Watchers

Forks

Releases20

Packages0

Uh oh!

Contributors12

Uh oh!

Languages

Packages