python 爬取免費簡歷模板網(wǎng)站的示例
代碼
# 免費的簡歷模板進行爬取本地保存 # http://sc.chinaz.com/jianli/free.html# http://sc.chinaz.com/jianli/free_2.htmlimport requestsfrom lxml import etreeimport osdirName = ’./resumeLibs’if not os.path.exists(dirName): os.mkdir(dirName)headers = { ’User-Agent’:’Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/85.0.4183.83 Safari/537.36’}url = ’http://sc.chinaz.com/jianli/free_%d.html’for page in range(1,2): if page == 1: new_url = ’http://sc.chinaz.com/jianli/free.html’ else: new_url = format(url%page) page_text = requests.get(url=new_url,headers=headers).text tree = etree.HTML(page_text) a_list = tree.xpath(’//div[@id='container']/div/p/a’) for a in a_list: a_src = a.xpath(’./@href’)[0] a_title = a.xpath(’./text()’)[0] a_title = a_title.encode(’iso-8859-1’).decode(’utf-8’) # 爬取下載頁面 page_text = requests.get(url=a_src,headers=headers).text tree = etree.HTML(page_text) dl_src = tree.xpath(’//div[@id='down']/div[2]/ul/li[8]/a/@href’)[0]resume_data = requests.get(url=dl_src,headers=headers).content resume_name = a_title resume_path = dirName + ’/’ + resume_name + ’.rar’ with open(resume_path,’wb’) as fp: fp.write(resume_data) print(resume_name,’下載成功!’)
爬取結(jié)果
以上就是python 爬取免費簡歷模板網(wǎng)站的示例的詳細內(nèi)容,更多關(guān)于python 爬取網(wǎng)站的資料請關(guān)注好吧啦網(wǎng)其它相關(guān)文章!
相關(guān)文章:
1. ASP 信息提示函數(shù)并作返回或者轉(zhuǎn)向2. windows服務器使用IIS時thinkphp搜索中文無效問題3. PHP設(shè)計模式中工廠模式深入詳解4. 淺談python出錯時traceback的解讀5. .NET中l(wèi)ambda表達式合并問題及解決方法6. Python importlib動態(tài)導入模塊實現(xiàn)代碼7. python matplotlib:plt.scatter() 大小和顏色參數(shù)詳解8. Ajax實現(xiàn)表格中信息不刷新頁面進行更新數(shù)據(jù)9. 利用promise及參數(shù)解構(gòu)封裝ajax請求的方法10. JSP數(shù)據(jù)交互實現(xiàn)過程解析
