python爬蟲添加請求頭代碼實例

2024-09-09 19:03:29

字體：大中小

來源：轉載

供稿：網友

這篇文章主要介紹了python爬蟲添加請求頭代碼實例,文中通過示例代碼介紹的非常詳細，對大家的學習或者工作具有一定的參考學習價值,需要的朋友可以參考下

request

import requestsheaders = {  # 'Accept': 'application/json, text/javascript, */*; q=0.01',  # 'Accept': '*/*',  # 'Accept-Language': 'zh-CN,zh;q=0.9,en;q=0.8,en-US;q=0.7',  # 'Cache-Control': 'no-cache',  # 'accept-encoding': 'gzip, deflate, br',  'User-Agent': 'Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/78.0.3904.97 Safari/537.36',  'Referer': 'https://www.google.com/'}resp = requests.get('http://httpbin.org/get', headers=headers)print(resp.content)

urllib

import urllib, urllib2def get_page_source(url):  headers = {'Accept': '*/*',        'Accept-Language': 'en-US,en;q=0.8',        'Cache-Control': 'max-age=0',        'User-Agent': 'Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/48.0.2564.116 Safari/537.36',        'Connection': 'keep-alive',        'Referer': 'http://www.baidu.com/'        }  req = urllib2.Request(url, None, headers)  response = urllib2.urlopen(req)  page_source = response.read()  return page_source

phantomjs請求頁面

from selenium import webdriverfrom selenium.webdriver.common.desired_capabilities import DesiredCapabilitiesdef get_headers_driver():  desire = DesiredCapabilities.PHANTOMJS.copy()  headers = {'Accept': '*/*',        'Accept-Language': 'en-US,en;q=0.8',        'Cache-Control': 'max-age=0',        'User-Agent': 'Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/48.0.2564.116 Safari/537.36',        'Connection': 'keep-alive',        'Referer': 'http://www.baidu.com/'        }  for key, value in headers.iteritems():    desire['phantomjs.page.customHeaders.{}'.format(key)] = value  driver = webdriver.PhantomJS(desired_capabilities=desire, service_args=['--load-images=yes'])#將yes改成no可以讓瀏覽器不加載圖片  return driver

以上就是本文的全部內容，希望對大家的學習有所幫助，也希望大家多多支持武林網之家。

上一篇：Python函數的返回值、匿名函數lambda、filter函數、map函數、red

下一篇：Python魔法方法容器部方法詳解