"illegal multibyte sequence" error from BeautifulSoup when Python 3

Question

.html saved to local disk, and I am using BeautifulSoup (bs4) to parse it.

It worked all fine until lately it's changed to Python 3.

I tested the same .html file in another machine Python 2, it works and returned the page contents.

soup = BeautifulSoup(open('page.html'), "lxml")

Machine with Python 3 doesn't work, and it says:

UnicodeDecodeError: 'gbk' codec can't decode byte 0x92 in position 298670: illegal multibyte sequence

Searched around and I tried below but neither worked: (be it 'r', or 'rb' doesn't make big difference)

soup = BeautifulSoup(open('page.html', 'r'), "lxml")
soup = BeautifulSoup(open('page.html', 'r'), 'html.parser')
soup = BeautifulSoup(open('page.html', 'r'), 'html5lib')
soup = BeautifulSoup(open('page.html', 'r'), 'xml')

How can I use Python 3 to parse this html page?

Thank you.

Sounds like the HTML is probably declaring the wrong encoding. I don't know how you'd override that, though. — user2357112
– user2357112, Commented Oct 9, 2019 at 8:38
When you say open('page.html', 'r'), then Python reads the document as plain-text and tries to decode it with some locale-dependent default, which is apparently GBK in your case. lxml should be fine with a binary stream however, so you should try opening it with open('page.html', 'rb'). Or you specify the correct encoding with the encoding= parameter. Note: depending on how the page was saved, the encoding declaration in the document may or may not be correct. — lenz
– lenz, Commented Oct 9, 2019 at 8:48
@lenz, it says "TypeError: 'from_encoding' is an invalid keyword argument for open()" — Mark K
– Mark K, Commented Oct 9, 2019 at 8:57
@lenz, it says "ValueError: binary mode doesn't take an encoding argument". — Mark K
– Mark K, Commented Oct 9, 2019 at 9:06

GPhilo · Accepted Answer · 2019-10-09 08:50:21Z

2

It worked all fine until lately it's changed to Python 3.

Python 3 has by default strings encoded in unicode, so when you open a file as text it will try to decode it. Python 2, on the other hand, uses bytestrings, instead and just returns the content of the file as-is. Try opening page.html as a byte object (open('page.html', 'rb')) and see if that works for you.

answered Oct 9, 2019 at 8:50

GPhilo

19.3k9 gold badges70 silver badges91 bronze badges

Sign up to request clarification or add additional context in comments.

5 Comments

Mark K Over a year ago

thanks for the reply. It give 1 more warning, says: UserWarning: No parser was explicitly specified, so I'm using the best available HTML parser for this system ("lxml"). This usually isn't a problem, but if you run this code on another system, or in a different virtual environment, it may use a different parser and behave differently.

GPhilo Over a year ago

That's a warning from BeautifulSoup, see here for how to get rid of it: stackoverflow.com/questions/33511544/…

Mark K Over a year ago

it's an additional warning message. The problem is still there.

lenz Over a year ago

@MarkK are you saying you opened the document in binary mode (open(..., 'rb')), and you still get a UnicodeDecodeError?

Mark K Over a year ago

@GPhilo, it seems the problem wasn't in the BeautifulSoup part. I posted some changes, which helped solved the problem.

Mark K · Accepted Answer · 2019-10-16 08:37:18Z

1

2 changes I done and not sure which one (or both) took the effect.

The computer was formatted and reinstalled so some settings are different.

1.In the language settings,

Administrative language settings > Change system locale >

Tick the box

Beta: Use Unicode UTF-8 for worldwide language support

2.on the coding, for example, this is the original line:

print (soup.find_all('span', attrs={'class': 'listing-row__price'})[0].text.strip().encode("utf-8"))

When the part ".encode("utf-8")" was removed, it worked.

update on 16th Oct. 2019 Above change works, but when the box is ticked. Fonts and texts in foreign language software doesn't display properly.
```
Beta: Use Unicode UTF-8 for worldwide language support
```

When the box was unticked, Fonts and texts in foreign language software are displayed well. But, problem in the question remains.

Solution with the box unticked - both foreign language software and Python codes work:

soup = BeautifulSoup(open(pages, 'r', encoding = 'utf-8', errors='ignore'), "lxml")

edited Oct 16, 2019 at 8:37

answered Oct 10, 2019 at 9:33

Mark K

9,50615 gold badges70 silver badges133 bronze badges

2 Comments

GPhilo Over a year ago

The second is the one that "solves" your problem, by simply printing the raw byte string instead of trying to encode it as UTF-8. You still have invalid unicode characters in your text, but if that's not important for your usage, ignoring them is a good option ;)

Mark K Over a year ago

@GPhilo, however it seems not - when the box "Beta: Use Unicode UTF-8 for worldwide language support" unticked, the problem pops again. (when the When the part ".encode("utf-8")" was removed, it doesn't worked.)

Collectives™ on Stack Overflow

"illegal multibyte sequence" error from BeautifulSoup when Python 3

2 Answers 2

5 Comments

2 Comments

Your Answer

Linked

Hot Network Questions

Collectives™ on Stack Overflow

2 Answers 2

5 Comments

2 Comments

Your Answer

Sign up or log in

Post as a guest

Linked

Related