# Capturing Webpages, including subpages

**URL:** https://discourse.devontechnologies.com/t/capturing-webpages-including-subpages/70897
**Category:** DEVONthink
**Created:** [May 6, 2022, 9:49am UTC](https://discourse.devontechnologies.com/t/capturing-webpages-including-subpages/70897 "2022-05-06T09:49:15Z")
**Posts on this page:** 20
**Page:** 1

<div class="post-metadata">

### Author: ![Wolkenhauer](https://discourse.devontechnologies.com/user_avatar/discourse.devontechnologies.com/wolkenhauer/32/10482_2.png) [@Wolkenhauer](https://discourse.devontechnologies.com/u/Wolkenhauer)
#### Post date: [May 6, 2022, 9:49am UTC](https://discourse.devontechnologies.com/t/capturing-webpages-including-subpages/70897/1 "2022-05-06T09:49:16Z")

</div>

What I like to do: Capturing a web-page, including subpages, for offline reading.

Hi Everyone

I checked the forum for a good description, but could not get it. The Help provides a section on Web Archive, which however does not download subpages.

The Download Manager is the way to go … but somehow I do not get it either … hence my call for your help 🙂

So, going to \> Window \> Download Manager, I have to windows with options to consider:

 ![2022-05-06_11-12-44](https://devontech-discourse.s3.dualstack.us-east-1.amazonaws.com/uploads/original/3X/7/f/7f4685d72071913a8ab9e78c951b68c8843e7c00.jpeg)  
 ![2022-05-06_11-13-13](https://devontech-discourse.s3.dualstack.us-east-1.amazonaws.com/uploads/original/3X/7/1/719b45dd1d766c3067acace1d7be1ab83c54a541.png)

Going for “Offline Archive” is the same as a Web Archive, not downloading subpages.

I then (regardless of the option chosen, and repeatedly) get an error message for the Global Inbox: “Failed database verification, please repair the database”.

 ![2022-05-06_11-25-53](https://devontech-discourse.s3.dualstack.us-east-1.amazonaws.com/uploads/original/3X/6/8/68fe6645b7034bde12f2e9466dd15a608f3b9928.png)

This finds one inconsistency and is easily repaired.

I then have inside a download folder in the global inbox, a folder and an html document of the main webpage that I like to read offline.

The folders show that all subpages are there but the main page, from which I want to go to subpages, that is still requiring the Internet. What I was expecting was a page where the links to subpages are now leading to the offline download. I imaged an offline copy of a webpage, so that I still go to that main page and click links to subpages.

Did I miss something?

Thank you in advance!

Olaf

---

<div class="post-metadata">

### Author: ![chrillek](https://discourse.devontechnologies.com/user_avatar/discourse.devontechnologies.com/chrillek/32/29340_2.png) [@chrillek](https://discourse.devontechnologies.com/u/chrillek)
#### Post date: [May 6, 2022, 10:11am UTC](https://discourse.devontechnologies.com/t/capturing-webpages-including-subpages/70897/2 "2022-05-06T10:11:35Z")

</div>

> [@Wolkenhauer](#):
>
> Capturing a web-page, including subpages,

What is the subpage of a web page? If you mean “I want to download all documents referred to by links in the original document”, I’d suggest you go for a `wget`, `curl` or such tools.

As an aside: The topic of capturing web documents has been discussed many times here. There’s no silver bullet to it. While webarchive might work in some cases, it has been (kind of) deprecated by Apple and is not supported on any other platform anyway. To reiterate my point: What you seemingly want (though I might have misunderstood your intentions) can be achieved with commandline download tools in a reliable and a very flexible, too. And you can of course script those with `do shell command` (AppleScript) or `doShellCommand` (JavaScript).

Also, capturing a page that is generated at least partly by JavaScript stored on the server (as is for example Apple’s developer documentation) will _always_ require an internet connection to display: The content is provided by the server when the page is loaded by the browser.

---

<div class="post-metadata">

### Author: ![Wolkenhauer](https://discourse.devontechnologies.com/user_avatar/discourse.devontechnologies.com/wolkenhauer/32/10482_2.png) [@Wolkenhauer](https://discourse.devontechnologies.com/u/Wolkenhauer)
#### Post date: [May 6, 2022, 10:48am UTC](https://discourse.devontechnologies.com/t/capturing-webpages-including-subpages/70897/3 "2022-05-06T10:48:10Z")

</div>

What I mean is a simple webpage with links, to other pages (no further levels). The download manager pulls all pages down, so this works well.

I was only hoping for the main page, with the link I used for the download, to then have its links to the subfolder/s that were downloaded, so that I can read the webpage, and its subpages, offline.

I read everything I could find on the download manager here before writing.

I am an ordinary user, no experience with scripts and was therefore looking for a solution within DT. No problem if it isn’t possible, I am not complaining. The result is already very close, I have all the files there, only the entry/main page with links to the folders is missing. I understand that there are complex scenarios with Java Script etc but my case was a simple two-layer scenario.

Thanks

Olaf

---

<div class="post-metadata">

### Author: ![rmschne](https://discourse.devontechnologies.com/letter_avatar_proxy/v4/letter/r/8baadc/32.png) [@rmschne](https://discourse.devontechnologies.com/u/rmschne)
#### Post date: [May 6, 2022, 10:59am UTC](https://discourse.devontechnologies.com/t/capturing-webpages-including-subpages/70897/4 "2022-05-06T10:59:55Z")

</div>

As suggested by @chrillek, check out and use wget and/or curl to do this.

---

<div class="post-metadata">

### Author: ![cgrunenberg](https://discourse.devontechnologies.com/user_avatar/discourse.devontechnologies.com/cgrunenberg/32/7172_2.png) [@cgrunenberg](https://discourse.devontechnologies.com/u/cgrunenberg)
#### Post date: [May 6, 2022, 11:03am UTC](https://discourse.devontechnologies.com/t/capturing-webpages-including-subpages/70897/5 "2022-05-06T11:03:54Z")

</div>

> [@Wolkenhauer](#):
>
> This finds one inconsistency and is easily repaired.

Was additional information logged to the _Log_ panel after verifying & repairing the database? Are you able to reproduce the issue?

---

<div class="post-metadata">

### Author: ![chrillek](https://discourse.devontechnologies.com/user_avatar/discourse.devontechnologies.com/chrillek/32/29340_2.png) [@chrillek](https://discourse.devontechnologies.com/u/chrillek)
#### Post date: [May 6, 2022, 11:04am UTC](https://discourse.devontechnologies.com/t/capturing-webpages-including-subpages/70897/6 "2022-05-06T11:04:23Z")

</div>

> [@Wolkenhauer](#):
>
> I have all the files there, only the entry/main page with links to the folders is missing

Well, that’s what the tools I mentioned are handling correctly (or at least `wget` does, I haven’t yet used `curl` for that).

> [@Wolkenhauer](#):
>
> I am an ordinary user, no experience with scripts

That shouldn’t stop you from _trying_ (which then leads to experience 😉). There are tons of explanations on how to use `curl` and `wegt` available online. It’s not rocket science nor will it break your computer (although you might fill up your hard disk if you do not limit the search depth).

---

<div class="post-metadata">

### Author: ![mbbntu](https://discourse.devontechnologies.com/letter_avatar_proxy/v4/letter/m/439d5e/32.png) [@mbbntu](https://discourse.devontechnologies.com/u/mbbntu)
#### Post date: [May 6, 2022, 11:19am UTC](https://discourse.devontechnologies.com/t/capturing-webpages-including-subpages/70897/7 "2022-05-06T11:19:16Z")

</div>

Some years ago I used a utility called SiteSucker that (it seems) will still work on MacOS. It might be worth investigating:

[https://ricks-apps.com/osx/sitesucker/index.html](https://ricks-apps.com/osx/sitesucker/index.html)

I don’t remember much about it, except that it worked for what I wanted to do at the time!

---

<div class="post-metadata">

### Author: ![BLUEFROG](https://discourse.devontechnologies.com/user_avatar/discourse.devontechnologies.com/bluefrog/32/135_2.png) [@BLUEFROG](https://discourse.devontechnologies.com/u/BLUEFROG)
#### Post date: [May 6, 2022, 12:57pm UTC](https://discourse.devontechnologies.com/t/capturing-webpages-including-subpages/70897/8 "2022-05-06T12:57:28Z")

</div>

Sheesh! I haven’t though about SiteSucker in years 😳😊

---

<div class="post-metadata">

### Author: ![mbbntu](https://discourse.devontechnologies.com/letter_avatar_proxy/v4/letter/m/439d5e/32.png) [@mbbntu](https://discourse.devontechnologies.com/u/mbbntu)
#### Post date: [May 6, 2022, 1:18pm UTC](https://discourse.devontechnologies.com/t/capturing-webpages-including-subpages/70897/9 "2022-05-06T13:18:16Z")

</div>

I’m getting really old 🙂

---

<div class="post-metadata">

### Author: ![Wolkenhauer](https://discourse.devontechnologies.com/user_avatar/discourse.devontechnologies.com/wolkenhauer/32/10482_2.png) [@Wolkenhauer](https://discourse.devontechnologies.com/u/Wolkenhauer)
#### Post date: [May 6, 2022, 1:18pm UTC](https://discourse.devontechnologies.com/t/capturing-webpages-including-subpages/70897/10 "2022-05-06T13:18:30Z")

</div>

Yes, I get the message for repair every time. I added a screenshot of the log window already above. Let me know if I can try anything else to resolve this.

---

<div class="post-metadata">

### Author: ![BLUEFROG](https://discourse.devontechnologies.com/user_avatar/discourse.devontechnologies.com/bluefrog/32/135_2.png) [@BLUEFROG](https://discourse.devontechnologies.com/u/BLUEFROG)
#### Post date: [May 6, 2022, 1:22pm UTC](https://discourse.devontechnologies.com/t/capturing-webpages-including-subpages/70897/11 "2022-05-06T13:22:02Z")

</div>

You said it was easily repaired. What did you repair?

---

<div class="post-metadata">

### Author: ![Wolkenhauer](https://discourse.devontechnologies.com/user_avatar/discourse.devontechnologies.com/wolkenhauer/32/10482_2.png) [@Wolkenhauer](https://discourse.devontechnologies.com/u/Wolkenhauer)
#### Post date: [May 6, 2022, 1:24pm UTC](https://discourse.devontechnologies.com/t/capturing-webpages-including-subpages/70897/12 "2022-05-06T13:24:11Z")

</div>

I understand that there are other tools, and thanks for the tips. I use DT primarily, every day, all the time, to gather material from the web, and also use web archives frequently. DT covers virtually all of my needs to gather material. So, when it came to having just the second layer of a simple webpage archived, I naturally hoped that it would do it as well. In a way, it does - the files are there, just the offline reading is not as convenient by having the main page linking to those downloaded subpages.

---

<div class="post-metadata">

### Author: ![Wolkenhauer](https://discourse.devontechnologies.com/user_avatar/discourse.devontechnologies.com/wolkenhauer/32/10482_2.png) [@Wolkenhauer](https://discourse.devontechnologies.com/u/Wolkenhauer)
#### Post date: [May 6, 2022, 1:25pm UTC](https://discourse.devontechnologies.com/t/capturing-webpages-including-subpages/70897/13 "2022-05-06T13:25:18Z")

</div>

The message says to repair the database (global inbox). I do that and than all is fine. It always finds one inconsistency after a download.

---

<div class="post-metadata">

### Author: ![Wolkenhauer](https://discourse.devontechnologies.com/user_avatar/discourse.devontechnologies.com/wolkenhauer/32/10482_2.png) [@Wolkenhauer](https://discourse.devontechnologies.com/u/Wolkenhauer)
#### Post date: [May 6, 2022, 1:28pm UTC](https://discourse.devontechnologies.com/t/capturing-webpages-including-subpages/70897/14 "2022-05-06T13:28:27Z")

</div>

We can have a Zoom session to look at it. I can send a Zoom link to quickly meet up and reproduce the scenario.

---

<div class="post-metadata">

### Author: ![cgrunenberg](https://discourse.devontechnologies.com/user_avatar/discourse.devontechnologies.com/cgrunenberg/32/7172_2.png) [@cgrunenberg](https://discourse.devontechnologies.com/u/cgrunenberg)
#### Post date: [May 6, 2022, 1:31pm UTC](https://discourse.devontechnologies.com/t/capturing-webpages-including-subpages/70897/15 "2022-05-06T13:31:23Z")

</div>

Please choose Help \> Report Bug while pressing the Alt modifier key and send the result to cgrunenberg - at - [devon-technologies.com](http://devon-technologies.com). This should be sufficient, thanks!

---

<div class="post-metadata">

### Author: ![Wolkenhauer](https://discourse.devontechnologies.com/user_avatar/discourse.devontechnologies.com/wolkenhauer/32/10482_2.png) [@Wolkenhauer](https://discourse.devontechnologies.com/u/Wolkenhauer)
#### Post date: [May 6, 2022, 1:50pm UTC](https://discourse.devontechnologies.com/t/capturing-webpages-including-subpages/70897/16 "2022-05-06T13:50:00Z")

</div>

Done

---

<div class="post-metadata">

### Author: ![cgrunenberg](https://discourse.devontechnologies.com/user_avatar/discourse.devontechnologies.com/cgrunenberg/32/7172_2.png) [@cgrunenberg](https://discourse.devontechnologies.com/u/cgrunenberg)
#### Post date: [May 6, 2022, 1:54pm UTC](https://discourse.devontechnologies.com/t/capturing-webpages-including-subpages/70897/17 "2022-05-06T13:54:10Z")

</div>

Thank you for the logs! Are there still any copies of the _Downloads_ group in the trash? That’s most likely causing the issue.

---

<div class="post-metadata">

### Author: ![Wolkenhauer](https://discourse.devontechnologies.com/user_avatar/discourse.devontechnologies.com/wolkenhauer/32/10482_2.png) [@Wolkenhauer](https://discourse.devontechnologies.com/u/Wolkenhauer)
#### Post date: [May 6, 2022, 2:25pm UTC](https://discourse.devontechnologies.com/t/capturing-webpages-including-subpages/70897/18 "2022-05-06T14:25:29Z")

</div>

Yes, during the various attempts, trying different options, I delete things, moving them to the trash. I take it, that emptying the trash is a good idea then 🙂

---

<div class="post-metadata">

### Author: ![BLUEFROG](https://discourse.devontechnologies.com/user_avatar/discourse.devontechnologies.com/bluefrog/32/135_2.png) [@BLUEFROG](https://discourse.devontechnologies.com/u/BLUEFROG)
#### Post date: [May 6, 2022, 3:11pm UTC](https://discourse.devontechnologies.com/t/capturing-webpages-including-subpages/70897/19 "2022-05-06T15:11:14Z")

</div>

> [@Wolkenhauer](#):
>
> I take it, that emptying the trash is a good idea then 🙂

Just like the waste bin in your kitchen, routinely emptying the trash in your databases is advisable.

---

<div class="post-metadata">

### Author: ![robbchadwick](https://discourse.devontechnologies.com/letter_avatar_proxy/v4/letter/r/e47774/32.png) [@robbchadwick](https://discourse.devontechnologies.com/u/robbchadwick)
#### Post date: [May 9, 2022, 9:08pm UTC](https://discourse.devontechnologies.com/t/capturing-webpages-including-subpages/70897/20 "2022-05-09T21:08:26Z")

</div>

SiteSucker is still going strong. The regular version is available in the Mac App Store — and there’s a Pro version available from the developer. I’ve always had good luck with it.

[Next page](https://discourse.devontechnologies.com/t/capturing-webpages-including-subpages/70897.md?page=2)
