How to use requests and Beautiful Soup to scrape a website that uses javascript? [closed]

Mangs

0人浏览 · 2022-08-25 01:23:03

Mangs · 2022-08-25 01:23:03 发布

Answer a question

I need to scrape this website:

https://sec.report/Ticker/AAPL

I need to get the CIK number 0000320193

When I do soup.prettify, it just says it needs to use javascript. Also, I don't want to open a web browser because it needs to be automated

I need to use the python beautiful soup and requests library

Answers

To get correct response from server, set correct User-Agent HTTP header:

import requests
from bs4 import BeautifulSoup


url = 'https://sec.report/Ticker/AAPL'
headers = {'User-Agent': 'Mozilla/5.0 (X11; Ubuntu; Linux x86_64; rv:78.0) Gecko/20100101 Firefox/78.0'}

soup = BeautifulSoup(requests.get(url, headers=headers).content, 'html.parser')
print(soup.h2.text) # or print(soup.h2.text.split()[-1]) for "0000320193"

Prints:

SEC CIK 0000320193

Python

Python社区为您提供最前沿的新闻资讯和知识内容

更多推荐

求助！为什么用InsCode部署会出现无限重定向？

Python

如何重塑熊猫。系列

问题:如何重塑熊猫。系列在我看来,它就像 pandas.Series 中的一个错误。 a = pd.Series([1,2,3,4]) b = a.reshape(2,2) b b 有类型 Series 但无法显示,最后一条语句给出异常,非常冗长,最后一行是“TypeError: %d format: a number is required, not numpy.ndarray”。 b.sha

Python

在哪里可以找到有关 Keras 中默认权重初始化器的文档? [复制]

问题:在哪里可以找到有关 Keras 中默认权重初始化器的文档? [复制] 我刚刚在这里](https://keras.io/initializers/)中阅读了有关[中的 Keras 权重初始化器的信息。在文档中,只介绍了不同的初始化程序。如: model.add(Dense(64, kernel_initializer='random_normal')) 当我没有指定kernel_initia