亚洲免费在线视频-亚洲啊v-久久免费精品视频-国产精品va-看片地址-成人在线视频网

您的位置:首頁技術文章
文章詳情頁

文本處理 - 求教使用python庫提取pdf的方法?

瀏覽:141日期:2022-09-02 11:18:15

問題描述

使用過pypdf 對英文pdf文檔處理比較簡單,但是對中文的支持好像不太好

使用過textract 看文檔支持的格式比較多方法也比較簡單,但是老師出錯

-- coding: utf-8 --

import textractimport pyPdfimport pdf2textimport pdfminerimport chardet

text = textract.process('F:ll.pdf',method = ’pdfminer’)print text

這個 出錯是編碼問題-- coding: utf-8 --

import textractimport pyPdfimport pdfminerimport chardet

text = textract.process('F:ll.pdf',method = ’pdfminer’)print text

這個出錯類型不清楚

少使用了pdf2text庫,但是出錯情況好像不一樣。

pdfminer庫還沒看過,看著好像麻煩一些, 求解一下解析提取中文的pdf的方法。謝謝

問題解答

回答1:

之前用過的pdfminer pip install pdfminer

# -*- coding: utf-8 -*-from bs4 import BeautifulSoupimport requestsimport refrom pdfminer.pdfinterp import PDFResourceManager, PDFPageInterpreterfrom pdfminer.converter import TextConverterfrom pdfminer.layout import LAParamsfrom cStringIO import StringIO#from io import StringIO for python3from io import openfrom pdfminer.pdfpage import PDFPagedef pdf_txt(url): rsrcmgr = PDFResourceManager() retstr = StringIO() codec = ’utf-8’ laparams = LAParams() device = TextConverter(rsrcmgr, retstr, codec=codec, laparams=laparams) f = requests.get(url).content fp = StringIO(f) interpreter = PDFPageInterpreter(rsrcmgr, device) password = '' maxpages = 0 caching = True pagenos = set() for page in PDFPage.get_pages(fp, pagenos, maxpages=maxpages, password=password, caching=caching, check_extractable=True):interpreter.process_page(page) fp.close() device.close() str = retstr.getvalue() retstr.close() return strtxt=tpdf_txt(’http://pythonscraping.com/pages/warandpeace/chapter1.pdf’)print txt#如果pdf含有中文,輸出到文件#open(’pdf.txt’,’wb’).write(txt)python readpdf.py’’’CHAPTER I'Well, Prince, so Genoa and Lucca are now just family estates oftheBuonapartes. But I warn you, if you don’t tell me that thismeans war,if you still try to defend the infamies and horrorsperpetrated bythat Antichrist- I really believe he is Antichrist- I willhavenothing more to do with you and you are no longer my friend,no longermy ’faithful slave,’ as you call yourself! But how do youdo? I seeI have frightened you- sit down and tell me all the news.'It was in July, 1805, and the speaker was the well-knownAnnaPavlovna Scherer, maid of honor and favorite of theEmpress MaryaFedorovna. With these words she greeted PrinceVasili Kuragin, a manof high rank and importance, who was thefirst to arrive at herreception. Anna Pavlovna had had a cough forsome days. She was, asshe said, suffering from la grippe; grippebeing then a new word inSt. Petersburg, used only by the elite.All her invitations without exception, written in French,anddelivered by a scarlet-liveried footman that morning, ran as’’’

標簽: Python 編程
相關文章:
主站蜘蛛池模板: 毛片成人永久免费视频 | 韩国一级黄色大片 | 99av在线| 手机看片国产欧美日韩高清 | 亚欧在线视频 | 亚洲国产品综合人成综合网站 | 成年人视频在线免费 | www.久草视频| 完整日本特级毛片 | 久久性感视频 | 99久久精品自在自看国产 | 国产丶欧美丶日韩丶不卡影视 | 美女张开双腿让男人桶 | 亚洲国产高清人在线 | 成人看片黄a免费看视频 | 精品一区二区三区波多野结衣 | 国产亚洲福利精品一区二区 | 亚洲精品国自产拍影院 | 在线免费观看精品 | 日韩精品一区二区三区四区 | 日韩毛片欧美一级a | 亚洲日产综合欧美一区二区 | 中文字幕在线视频网站 | 狠狠色丁香九九婷婷综合五月 | 日本免费人成黄页在线观看视频 | 欧美精品综合一区二区三区 | 国产精品9999久久久久 | 在线日本视频 | 午夜爽爽视频 | 喷潮白浆直流在线播放 | 一区二区三区久久 | 欧美成人 一区二区三区 | 亚洲an日韩专区在线 | 亚洲精品98久久久久久中文字幕 | 中国一级毛片欧美一级毛片 | 怡红院免费的全部视频 | 麻豆视频一区 | 欧美日韩精品乱国产538 | 黄色三级网址 | 日日狠狠久久偷偷四色综合免费 | 国产日韩精品在线 |