Revisions to Reduce run time of NLP approximate matching code

Removed sentence

Source Link

edited May 20, 2020 at 13:55

33
6

The code below matches a list of features to a large corpus and returns the sub-query match with a score above 80. The challenge is the list of features on the full data-set is > 5,000 and comparing to multiple documents. Therefore is taking too long to work using the fuzzywuzzy package. I tried using the textdistance.cosine with minimal improvement in speed.

The code below matches a list of features to a large corpus and returns the sub-query match with a score above 80. The challenge is the list of features on the full data-set is > 5,000 and comparing to multiple documents. Therefore is taking too long to work using the fuzzywuzzy package.

Tweeted twitter.com/StackCodeReview/status/1262488160237486081

occurred May 18, 2020 at 21:00

fix code block

Source Link

edited Apr 28, 2020 at 5:14

greybeard

7.8k
3
21
56

from fuzzywuzzy import fuzz
import re
    
document = """If you're shopping within the Toyota family, the Highlander offers appreciably more space than the RAV4, both in terms of cargo capacity and its extra row of seats. It also has a deeper, more accessible space than what's in the 4Runner.
That said, the Highlander is one of the smallest three-row crossovers available. Apart from the Kia Sorento and maybe the Mazda CX-9, you're going to find more cargo capacity and passenger space in the Highlander's competitors. That's especially true in the third row. The second row slides a bit more to grant extra legroom now, but the third row remains awfully close to the floor, and it won't be long before your growing kids will feel cramped and claustrophobic in the way-back. Full-size teens and adults will be flat-out grumpy.
That said, the Highlander's smaller size might be just right for many buyers who appreciate its more manageable dimensions when parking or maneuvering in tight spots. Plus, if you only need that third row for occasional use and just a little more space than what a RAV4 provides, it really won't matter that the Highlander can't match its competitors' jumbo size.
We expect pricing for the 2020 Highlander to be announced closer to its on-sale date in December 2019, with the Hybrid arriving in February 2020. Specifically, it should correspond with our first test drive opportunity, likely in November. We do have a pretty comprehensive features breakdown, however, which you can see below.
Standard equipment on the Highlander L includes 18-inch alloy wheels, three-zone automatic climate control, accident avoidance tech features (see safety section below), full-speed adaptive cruise control, LED headlights, rear privacy glass, proximity entry and push-button start, an eight-way power driver seat and the 8-inch touchscreen. The LE additions include a power liftgate, blind-spot warning, LED foglamps, and a leather-wrapped steering wheel.
The XLE additions include automatic headlights, roof rails, a sunroof, heated front seats, driver power lumbar, a four-way power passenger seat, SofTex vinyl upholstery, second-row sunshades and an auto-dimming rearview mirror.
The Limited additions include 20-inch wheels, a handsfree power liftgate, upgraded LED headlights, a cargo cover, driver memory settings, ventilated front seats, leather upholstery, integrated navigation and a JBL sound system upgrade.
The Platinum additions include adaptive and self-leveling headlights, automatic wipers, a panoramic sunroof bird's-eye parking camera, a head-up display, a digital rearview mirror camera, perforated leather upholstery, heated second-row seats and a 12.3-inch touchscreen.
"""
features =["steering","touch screen","LED headlight"]

def findcarfeatures(features, document, match=80):
    result=[]
    for feature in features:
        lenfeature = len(feature.split(" "))
        word_tokens = nltk.word_tokenize(document)
        #filterd_word_tokens = [w for w in word_tokens if not w in stop_words]
        for i in range (len(word_tokens)-lenfeature+1):
            wordtocompare = ""
            j=0
            for j in range(i, i+lenfeature):
                if re.search(r'[,!?{}\[\]\"\"\'\']',word_tokens[j]):
                    break
                wordtocompare = wordtocompare+" "+word_tokens[j].lower()
            wordtocompare.strip()
            if not wordtocompare=="":
                if(fuzz.ratio(wordtocompare,feature.lower())> match):
                    result.append([wordtocompare,feature,i,j])
    return result

findcarfeatures(features,document)

Out[90]: 
[[' steering', 'steering', 353, 353],
 [' touchscreen .', 'touch screen', 334, 335],
 [' touchscreen .', 'touch screen', 474, 475],
 [' led headlights', 'LED headlight', 313, 314],
 [' headlights', 'LED headlight', 314, 315],
 [' headlights', 'LED headlight', 361, 362],
 [' led headlights', 'LED headlight', 408, 409],
 [' headlights', 'LED headlight', 409, 410],
 [' headlights', 'LED headlight', 442, 443]]
```

import pandas as pd
from fuzzywuzzy import fuzz
import re
    
document = """If you're shopping within the Toyota family, the Highlander offers appreciably more space than the RAV4, both in terms of cargo capacity and its extra row of seats. It also has a deeper, more accessible space than what's in the 4Runner.
That said, the Highlander is one of the smallest three-row crossovers available. Apart from the Kia Sorento and maybe the Mazda CX-9, you're going to find more cargo capacity and passenger space in the Highlander's competitors. That's especially true in the third row. The second row slides a bit more to grant extra legroom now, but the third row remains awfully close to the floor, and it won't be long before your growing kids will feel cramped and claustrophobic in the way-back. Full-size teens and adults will be flat-out grumpy.
That said, the Highlander's smaller size might be just right for many buyers who appreciate its more manageable dimensions when parking or maneuvering in tight spots. Plus, if you only need that third row for occasional use and just a little more space than what a RAV4 provides, it really won't matter that the Highlander can't match its competitors' jumbo size.
We expect pricing for the 2020 Highlander to be announced closer to its on-sale date in December 2019, with the Hybrid arriving in February 2020. Specifically, it should correspond with our first test drive opportunity, likely in November. We do have a pretty comprehensive features breakdown, however, which you can see below.
Standard equipment on the Highlander L includes 18-inch alloy wheels, three-zone automatic climate control, accident avoidance tech features (see safety section below), full-speed adaptive cruise control, LED headlights, rear privacy glass, proximity entry and push-button start, an eight-way power driver seat and the 8-inch touchscreen. The LE additions include a power liftgate, blind-spot warning, LED foglamps, and a leather-wrapped steering wheel.
The XLE additions include automatic headlights, roof rails, a sunroof, heated front seats, driver power lumbar, a four-way power passenger seat, SofTex vinyl upholstery, second-row sunshades and an auto-dimming rearview mirror.
The Limited additions include 20-inch wheels, a handsfree power liftgate, upgraded LED headlights, a cargo cover, driver memory settings, ventilated front seats, leather upholstery, integrated navigation and a JBL sound system upgrade.
The Platinum additions include adaptive and self-leveling headlights, automatic wipers, a panoramic sunroof bird's-eye parking camera, a head-up display, a digital rearview mirror camera, perforated leather upholstery, heated second-row seats and a 12.3-inch touchscreen.
"""
features =["steering","touch screen","LED headlight"]

def findcarfeatures(features, document, match=80):
    result=[]
    for feature in features:
        lenfeature = len(feature.split(" "))
        word_tokens = nltk.word_tokenize(document)
        #filterd_word_tokens = [w for w in word_tokens if not w in stop_words]
        for i in range (len(word_tokens)-lenfeature+1):
            wordtocompare = ""
            j=0
            for j in range(i, i+lenfeature):
                if re.search(r'[,!?{}\[\]\"\"\'\']',word_tokens[j]):
                    break
                wordtocompare = wordtocompare+" "+word_tokens[j].lower()
            wordtocompare.strip()
            if not wordtocompare=="":
                if(fuzz.ratio(wordtocompare,feature.lower())> match):
                    result.append([wordtocompare,feature,i,j])
    return result

findcarfeatures(features,document)

Out[90]: 
[[' steering', 'steering', 353, 353],
 [' touchscreen .', 'touch screen', 334, 335],
 [' touchscreen .', 'touch screen', 474, 475],
 [' led headlights', 'LED headlight', 313, 314],
 [' headlights', 'LED headlight', 314, 315],
 [' headlights', 'LED headlight', 361, 362],
 [' led headlights', 'LED headlight', 408, 409],
 [' headlights', 'LED headlight', 409, 410],
 [' headlights', 'LED headlight', 442, 443]]

from fuzzywuzzy import fuzz
import re
    
document = """If you're shopping within the Toyota family, the Highlander offers appreciably more space than the RAV4, both in terms of cargo capacity and its extra row of seats. It also has a deeper, more accessible space than what's in the 4Runner.
That said, the Highlander is one of the smallest three-row crossovers available. Apart from the Kia Sorento and maybe the Mazda CX-9, you're going to find more cargo capacity and passenger space in the Highlander's competitors. That's especially true in the third row. The second row slides a bit more to grant extra legroom now, but the third row remains awfully close to the floor, and it won't be long before your growing kids will feel cramped and claustrophobic in the way-back. Full-size teens and adults will be flat-out grumpy.
That said, the Highlander's smaller size might be just right for many buyers who appreciate its more manageable dimensions when parking or maneuvering in tight spots. Plus, if you only need that third row for occasional use and just a little more space than what a RAV4 provides, it really won't matter that the Highlander can't match its competitors' jumbo size.
We expect pricing for the 2020 Highlander to be announced closer to its on-sale date in December 2019, with the Hybrid arriving in February 2020. Specifically, it should correspond with our first test drive opportunity, likely in November. We do have a pretty comprehensive features breakdown, however, which you can see below.
Standard equipment on the Highlander L includes 18-inch alloy wheels, three-zone automatic climate control, accident avoidance tech features (see safety section below), full-speed adaptive cruise control, LED headlights, rear privacy glass, proximity entry and push-button start, an eight-way power driver seat and the 8-inch touchscreen. The LE additions include a power liftgate, blind-spot warning, LED foglamps, and a leather-wrapped steering wheel.
The XLE additions include automatic headlights, roof rails, a sunroof, heated front seats, driver power lumbar, a four-way power passenger seat, SofTex vinyl upholstery, second-row sunshades and an auto-dimming rearview mirror.
The Limited additions include 20-inch wheels, a handsfree power liftgate, upgraded LED headlights, a cargo cover, driver memory settings, ventilated front seats, leather upholstery, integrated navigation and a JBL sound system upgrade.
The Platinum additions include adaptive and self-leveling headlights, automatic wipers, a panoramic sunroof bird's-eye parking camera, a head-up display, a digital rearview mirror camera, perforated leather upholstery, heated second-row seats and a 12.3-inch touchscreen.
"""
features =["steering","touch screen","LED headlight"]

def findcarfeatures(features, document, match=80):
    result=[]
    for feature in features:
        lenfeature = len(feature.split(" "))
        word_tokens = nltk.word_tokenize(document)
        #filterd_word_tokens = [w for w in word_tokens if not w in stop_words]
        for i in range (len(word_tokens)-lenfeature+1):
            wordtocompare = ""
            j=0
            for j in range(i, i+lenfeature):
                if re.search(r'[,!?{}\[\]\"\"\'\']',word_tokens[j]):
                    break
                wordtocompare = wordtocompare+" "+word_tokens[j].lower()
            wordtocompare.strip()
            if not wordtocompare=="":
                if(fuzz.ratio(wordtocompare,feature.lower())> match):
                    result.append([wordtocompare,feature,i,j])
    return result

findcarfeatures(features,document)

Out[90]: 
[[' steering', 'steering', 353, 353],
 [' touchscreen .', 'touch screen', 334, 335],
 [' touchscreen .', 'touch screen', 474, 475],
 [' led headlights', 'LED headlight', 313, 314],
 [' headlights', 'LED headlight', 314, 315],
 [' headlights', 'LED headlight', 361, 362],
 [' led headlights', 'LED headlight', 408, 409],
 [' headlights', 'LED headlight', 409, 410],
 [' headlights', 'LED headlight', 442, 443]]
```

import pandas as pd
from fuzzywuzzy import fuzz
import re
    
document = """If you're shopping within the Toyota family, the Highlander offers appreciably more space than the RAV4, both in terms of cargo capacity and its extra row of seats. It also has a deeper, more accessible space than what's in the 4Runner.
That said, the Highlander is one of the smallest three-row crossovers available. Apart from the Kia Sorento and maybe the Mazda CX-9, you're going to find more cargo capacity and passenger space in the Highlander's competitors. That's especially true in the third row. The second row slides a bit more to grant extra legroom now, but the third row remains awfully close to the floor, and it won't be long before your growing kids will feel cramped and claustrophobic in the way-back. Full-size teens and adults will be flat-out grumpy.
That said, the Highlander's smaller size might be just right for many buyers who appreciate its more manageable dimensions when parking or maneuvering in tight spots. Plus, if you only need that third row for occasional use and just a little more space than what a RAV4 provides, it really won't matter that the Highlander can't match its competitors' jumbo size.
We expect pricing for the 2020 Highlander to be announced closer to its on-sale date in December 2019, with the Hybrid arriving in February 2020. Specifically, it should correspond with our first test drive opportunity, likely in November. We do have a pretty comprehensive features breakdown, however, which you can see below.
Standard equipment on the Highlander L includes 18-inch alloy wheels, three-zone automatic climate control, accident avoidance tech features (see safety section below), full-speed adaptive cruise control, LED headlights, rear privacy glass, proximity entry and push-button start, an eight-way power driver seat and the 8-inch touchscreen. The LE additions include a power liftgate, blind-spot warning, LED foglamps, and a leather-wrapped steering wheel.
The XLE additions include automatic headlights, roof rails, a sunroof, heated front seats, driver power lumbar, a four-way power passenger seat, SofTex vinyl upholstery, second-row sunshades and an auto-dimming rearview mirror.
The Limited additions include 20-inch wheels, a handsfree power liftgate, upgraded LED headlights, a cargo cover, driver memory settings, ventilated front seats, leather upholstery, integrated navigation and a JBL sound system upgrade.
The Platinum additions include adaptive and self-leveling headlights, automatic wipers, a panoramic sunroof bird's-eye parking camera, a head-up display, a digital rearview mirror camera, perforated leather upholstery, heated second-row seats and a 12.3-inch touchscreen.
"""
features =["steering","touch screen","LED headlight"]

def findcarfeatures(features, document, match=80):
    result=[]
    for feature in features:
        lenfeature = len(feature.split(" "))
        word_tokens = nltk.word_tokenize(document)
        #filterd_word_tokens = [w for w in word_tokens if not w in stop_words]
        for i in range (len(word_tokens)-lenfeature+1):
            wordtocompare = ""
            j=0
            for j in range(i, i+lenfeature):
                if re.search(r'[,!?{}\[\]\"\"\'\']',word_tokens[j]):
                    break
                wordtocompare = wordtocompare+" "+word_tokens[j].lower()
            wordtocompare.strip()
            if not wordtocompare=="":
                if(fuzz.ratio(wordtocompare,feature.lower())> match):
                    result.append([wordtocompare,feature,i,j])
    return result

findcarfeatures(features,document)

Out[90]: 
[[' steering', 'steering', 353, 353],
 [' touchscreen .', 'touch screen', 334, 335],
 [' touchscreen .', 'touch screen', 474, 475],
 [' led headlights', 'LED headlight', 313, 314],
 [' headlights', 'LED headlight', 314, 315],
 [' headlights', 'LED headlight', 361, 362],
 [' led headlights', 'LED headlight', 408, 409],
 [' headlights', 'LED headlight', 409, 410],
 [' headlights', 'LED headlight', 442, 443]]

deleted 59 characters in body

Source Link

edited Apr 28, 2020 at 5:07

Jamal

35.2k
13
134
238

Per the Spyder profiler, the bottlenecks are at if(fuzz.ratio(wordtocompare,feature.lower())> match) and _find_and_load_unlocked:

if(fuzz.ratio(wordtocompare,feature.lower())> match) and _find_and_load_unlocked

Would vectorizing the code in its current form help or is there a faster approximate matcher that accounts for matching a sub-query (information extraction) of text compared to a defined list? Has anyone had success using polyleven and port the results back into Python? Any help would be greatly appreciated. Thank you for your time.

Per the Spyder profiler, the bottlenecks are at if(fuzz.ratio(wordtocompare,feature.lower())> match) and _find_and_load_unlocked

Would vectorizing the code in its current form help or is there a faster approximate matcher that accounts for matching a sub-query (information extraction) of text compared to a defined list? Has anyone had success using polyleven and port the results back into Python? Any help would be greatly appreciated. Thank you for your time.

Per the Spyder profiler, the bottlenecks are at:

if(fuzz.ratio(wordtocompare,feature.lower())> match) and _find_and_load_unlocked

Would vectorizing the code in its current form help or is there a faster approximate matcher that accounts for matching a sub-query (information extraction) of text compared to a defined list? Has anyone had success using polyleven and port the results back into Python?

added 25 characters in body

Source Link

edited Apr 28, 2020 at 4:16

Rtimeseries

33
6

Loading

added 18 characters in body

Source Link

edited Apr 28, 2020 at 3:57

Rtimeseries

33
6

Loading

edited tags

Link

edited Apr 27, 2020 at 20:38

Rtimeseries

33
6

Loading

added 60 characters in body; edited tags

Source Link

edited Apr 25, 2020 at 14:40

Rtimeseries

33
6

Loading

edited tags

Link

edited Apr 20, 2020 at 15:29

Peilonrayz ♦

44.6k
7
80
158

Loading

edited tags

Link

edited Apr 20, 2020 at 15:20

Rtimeseries

33
6

Loading

edited tags

Link

edited Apr 18, 2020 at 3:14

Rtimeseries

33
6

Loading

Rollback to Revision 4

Source Link

edited Apr 16, 2020 at 16:49

Sᴀᴍ Onᴇᴌᴀ ♦

29.6k
16
46
203

Loading

deleted 37 characters in body

Source Link

edited Apr 16, 2020 at 16:22

Rtimeseries

33
6

Loading

added 130 characters in body

Source Link

edited Apr 15, 2020 at 20:26

Rtimeseries

33
6

Loading

deleted 1 character in body

Source Link

edited Apr 15, 2020 at 20:03

Rtimeseries

33
6

Loading

deleted 1 character in body

Source Link

edited Apr 15, 2020 at 20:01

Rtimeseries

33
6

Loading

Source Link

asked Apr 15, 2020 at 19:52

Rtimeseries

33
6

Loading

Stack Exchange Network

Return to Question