# EgoTracks dataset download failure

**URL:** <https://discuss.ego4d-data.org/t/egotracks-dataset-download-failure/218>\
**Category:** Q&A\
**Created:** [March 4, 2023, 10:17pm UTC](https://discuss.ego4d-data.org/t/egotracks-dataset-download-failure/218 "2023-03-04T22:17:33Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![aram](https://avatars.discourse-cdn.com/v4/letter/a/a183cd/32.png) [@aram](https://discuss.ego4d-data.org/u/aram)\
**Post date:** [March 4, 2023, 10:17pm UTC](https://discuss.ego4d-data.org/t/egotracks-dataset-download-failure/218/1 "2023-03-04T22:17:33Z")

</div>

Hello.  
I have received my aws cli license from ego4d yesterday, and I’m trying to download the “EgoTracks” dataset.

I can successsfully download the viz and annotations but egotracks videos is failing.

`ego4d --output_directory="~/scratch/data/tracking/ego4d" --datasets egotracks`  
**output:**  
_botocore.exceptions.ClientError: An error occurred (403) when calling the HeadObject operation: Forbidden_

Also I wanted to ask how large (in GB) the EgoTracks dataset is? can only find it consists of 5.9K videos but I couldn’t find anywhere it states the actual size.

I have followed the instructions based on  
[egotracks download instructions](https://github.com/EGO4D/docs/blob/main/docs/data/egotracks.md)  
and  
[ego4d cli instructions](https://github.com/facebookresearch/Ego4d/blob/main/ego4d/cli/README.md)

---

<div class="post-metadata">

**Author:** ![vkvats](https://avatars.discourse-cdn.com/v4/letter/v/c4cdca/32.png) [@vkvats](https://discuss.ego4d-data.org/u/vkvats)\
**Post date:** [March 5, 2023, 4:37pm UTC](https://discuss.ego4d-data.org/t/egotracks-dataset-download-failure/218/2 "2023-03-05T16:37:27Z")

</div>

Hi

I am running into the same problem. If I do not use the region in config file then it throws ValueError: Invalid endpoint and if I use us-east-1 or us-east-2 then it throws the error stated above. Did you find any solution?

Thanks,

---

<div class="post-metadata">

**Author:** ![aram](https://avatars.discourse-cdn.com/v4/letter/a/a183cd/32.png) [@aram](https://discuss.ego4d-data.org/u/aram)\
**Post date:** [March 5, 2023, 11:00pm UTC](https://discuss.ego4d-data.org/t/egotracks-dataset-download-failure/218/3 "2023-03-05T23:00:49Z")

</div>

Hello I have not been able to resolve this. Also i think `--datasets annotations_540ss` also doesn’t work… Please let me know where i can find the downscaled 540ss annotations!

---

<div class="post-metadata">

**Author:** ![gene](https://avatars.discourse-cdn.com/v4/letter/g/43a26b/32.png) [@gene](https://discuss.ego4d-data.org/u/gene)\
**Post date:** [March 6, 2023, 7:45pm UTC](https://discuss.ego4d-data.org/t/egotracks-dataset-download-failure/218/4 "2023-03-06T19:45:48Z")

</div>

Please try with the `--version v2` argument as well. Can you confirm that works for EgoTracks?

(Looking at 540ss, will come back shortly.)

---

<div class="post-metadata">

**Author:** ![aram](https://avatars.discourse-cdn.com/v4/letter/a/a183cd/32.png) [@aram](https://discuss.ego4d-data.org/u/aram)\
**Post date:** [March 6, 2023, 8:18pm UTC](https://discuss.ego4d-data.org/t/egotracks-dataset-download-failure/218/5 "2023-03-06T20:18:13Z")

</div>

–version v2 argument works, but it says the dataset is only 2.4 GB which I am highly doubtful it is the size of the entire egotracks dataset?

`ego4d --output_directory ./ --datasets egotracks --version v2`  
`Expected size of downloaded files is 2.4 GB. Do you want to start the download? `

Could you confirm with me that 2.4GB is the correct size of the entire EgoTracks dataset?

---

<div class="post-metadata">

**Author:** ![aram](https://avatars.discourse-cdn.com/v4/letter/a/a183cd/32.png) [@aram](https://discuss.ego4d-data.org/u/aram)\
**Post date:** [March 6, 2023, 9:18pm UTC](https://discuss.ego4d-data.org/t/egotracks-dataset-download-failure/218/6 "2023-03-06T21:18:33Z")

</div>

I think the `--datasets egotracks` only gives the annotation json files.  
So do I have to download the video data by the following command???  
`ego4d --output_directory ./ --datasets full_scale --benchmark EM --version v2`

---

<div class="post-metadata">

**Author:** ![aram](https://avatars.discourse-cdn.com/v4/letter/a/a183cd/32.png) [@aram](https://discuss.ego4d-data.org/u/aram)\
**Post date:** [March 8, 2023, 7:16pm UTC](https://discuss.ego4d-data.org/t/egotracks-dataset-download-failure/218/7 "2023-03-08T19:16:01Z")

</div>

Hello. I have downloaded version 2 of the full-scale videos for EM Benchmark, but while running the preprocessing step, I meet this error.

`python tools/preprocess/extract_ego4d_clip_frames.py`

```bash
line 77, in extract_clip_ids
clip_uids.append(c["exported_clip_uid"])
KeyError: 'exported_clip_uid'

```

Could you please provide a full guide for the egotracks dataset download & preprocess so I can follow?  
Thank you.

---

<div class="post-metadata">

**Author:** ![haotang](https://avatars.discourse-cdn.com/v4/letter/h/9fc348/32.png) [@haotang](https://discuss.ego4d-data.org/u/haotang)\
**Post date:** [March 9, 2023, 1:28am UTC](https://discuss.ego4d-data.org/t/egotracks-dataset-download-failure/218/8 "2023-03-09T01:28:45Z")

</div>

Hi, which annotation\_path are you using (train, val or test)?

- This should not happen with the challenge test set, but if it does, please let us know!
- For the train and val, we are working on pushing an updated preprocess that should solve the problem. The workaround is: You can simple ignore these clips (should be less than 1%). We don’t have the exported\_clip\_uid for these frames because of conversion error.

---

<div class="post-metadata">

**Author:** ![aram](https://avatars.discourse-cdn.com/v4/letter/a/a183cd/32.png) [@aram](https://discuss.ego4d-data.org/u/aram)\
**Post date:** [March 9, 2023, 5:30pm UTC](https://discuss.ego4d-data.org/t/egotracks-dataset-download-failure/218/9 "2023-03-09T17:30:11Z")

</div>

Thank you for the reply.  
Can you please confirm for me that the download script i used is correct?

`ego4d --output_directory ./ --datasets egotracks full_scale --benchmark EM --version v2`

Total size was about 2.7 TB

---

<div class="post-metadata">

**Author:** ![haotang](https://avatars.discourse-cdn.com/v4/letter/h/9fc348/32.png) [@haotang](https://discuss.ego4d-data.org/u/haotang)\
**Post date:** [March 9, 2023, 8:11pm UTC](https://discuss.ego4d-data.org/t/egotracks-dataset-download-failure/218/10 "2023-03-09T20:11:35Z")

</div>

Hi, I am working on confirming the download script, but we don’t need full\_scale, only the clips for EM are needed. So it should be the following:

```auto
ego4d --output_directory ./ --datasets egotracks clips --benchmark EM --version v2

```

---

<div class="post-metadata">

**Author:** ![aram](https://avatars.discourse-cdn.com/v4/letter/a/a183cd/32.png) [@aram](https://discuss.ego4d-data.org/u/aram)\
**Post date:** [March 9, 2023, 8:49pm UTC](https://discuss.ego4d-data.org/t/egotracks-dataset-download-failure/218/11 "2023-03-09T20:49:14Z")

</div>

Thank you so much for the fast reply! I think I will try to re-download only the clip dataset in the meanwhile. Please let me know when the preprocess script is cleaned :). Thanks @haotang!

---

<div class="post-metadata">

**Author:** ![aram](https://avatars.discourse-cdn.com/v4/letter/a/a183cd/32.png) [@aram](https://discuss.ego4d-data.org/u/aram)\
**Post date:** [March 10, 2023, 6:50pm UTC](https://discuss.ego4d-data.org/t/egotracks-dataset-download-failure/218/12 "2023-03-10T18:50:12Z")

</div>

@haotang  
Just as an FYI. `Skipping 113 videos...Total 3433 to be processed ...` Is the number of videos that don’t have “exported\_clip\_uid” field in annotations

---

<div class="post-metadata">

**Author:** ![haotang](https://avatars.discourse-cdn.com/v4/letter/h/9fc348/32.png) [@haotang](https://discuss.ego4d-data.org/u/haotang)\
**Post date:** [March 10, 2023, 7:31pm UTC](https://discuss.ego4d-data.org/t/egotracks-dataset-download-failure/218/13 "2023-03-10T19:31:44Z")

</div>

Hi @aram Thanks for the sharing the numbers! This looks correct to me. I ignored a few more videos because of issues with the frame conversion for certain bounding boxes. I created a fix in [Egotracks fix by tanghaotommy · Pull Request #42 · EGO4D/episodic-memory · GitHub](https://github.com/EGO4D/episodic-memory/pull/42) and am waiting for review. But you may take a look and use that.

---

<div class="post-metadata">

**Author:** ![aram](https://avatars.discourse-cdn.com/v4/letter/a/a183cd/32.png) [@aram](https://discuss.ego4d-data.org/u/aram)\
**Post date:** [March 10, 2023, 8:25pm UTC](https://discuss.ego4d-data.org/t/egotracks-dataset-download-failure/218/14 "2023-03-10T20:25:38Z")

</div>

Thank you for making this adjustment!!

---

<div class="post-metadata">

**Author:** ![aram](https://avatars.discourse-cdn.com/v4/letter/a/a183cd/32.png) [@aram](https://discuss.ego4d-data.org/u/aram)\
**Post date:** [March 11, 2023, 7:22pm UTC](https://discuss.ego4d-data.org/t/egotracks-dataset-download-failure/218/15 "2023-03-11T19:22:15Z")

</div>

I have tried running the preprocess code but it hangs. Also I can no longer cd or ls into the drive storing the data. Currently I am using a 4TB ssd to store all the data. Could you please provide reference on what is the total disk space that is required to run the preprocess script?

---

<div class="post-metadata">

**Author:** ![haotang](https://avatars.discourse-cdn.com/v4/letter/h/9fc348/32.png) [@haotang](https://discuss.ego4d-data.org/u/haotang)\
**Post date:** [March 15, 2023, 12:09am UTC](https://discuss.ego4d-data.org/t/egotracks-dataset-download-failure/218/16 "2023-03-15T00:09:05Z")

</div>

I believe the clips themselves are less than 1TB. The preprocessed data takes about 800GB (only annotated frames).

---

<div class="post-metadata">

**Author:** ![aram](https://avatars.discourse-cdn.com/v4/letter/a/a183cd/32.png) [@aram](https://discuss.ego4d-data.org/u/aram)\
**Post date:** [March 15, 2023, 6:05am UTC](https://discuss.ego4d-data.org/t/egotracks-dataset-download-failure/218/17 "2023-03-15T06:05:54Z")

</div>

Thanks. I have a 4 TB ssd, and I have a problem where the extraction code hangs in the middle, and I cannot ls into my ssd. (probably the extraction process/thread is not exitting??)

So I have cancelled - restarted multiple times but know the extracted frames folder is ~ 3.8 TB. (disk space is basically full)  
Do you suspect anything going wrong?

Major problems

1. frame extraction process hanging (probably due to 2? but not sure)
2. Disk space requirement \> 8 TB

I did some calculation where  
each video ~ 8 min with 30fps = 8 \* 60 \* 30 = 14400 frames. (I checked and the extracted folder actually has 14400 frames)  
Each frame ~ 200 KB  
Each video image folder (extracted frames) = 14400 \* 200 KB = 2.88 GB.  
Train set includes 3000 videos which leads to 8.6 TB disk space for extracted frames.

Would really appreciate your reply!

---

<div class="post-metadata">

**Author:** ![aram](https://avatars.discourse-cdn.com/v4/letter/a/a183cd/32.png) [@aram](https://discuss.ego4d-data.org/u/aram)\
**Post date:** [March 15, 2023, 6:14am UTC](https://discuss.ego4d-data.org/t/egotracks-dataset-download-failure/218/18 "2023-03-15T06:14:38Z")

</div>

I noticed that you mentioned we should be extracting _“only annotated frames”_ but I guess the current preprocessing code is extracting all frames?

Please correct me if I am wrong.

---

<div class="post-metadata">

**Author:** ![haotang](https://avatars.discourse-cdn.com/v4/letter/h/9fc348/32.png) [@haotang](https://discuss.ego4d-data.org/u/haotang)\
**Post date:** [March 15, 2023, 3:25pm UTC](https://discuss.ego4d-data.org/t/egotracks-dataset-download-failure/218/19 "2023-03-15T15:25:01Z")

</div>

Yes, true. We annotate at 5FPS, so it should be fine if only extracting those frames. If you extract at 30 FPS, the disk space is not enough. Please take a look at the pull request here: [Egotracks fix by tanghaotommy · Pull Request #42 · EGO4D/episodic-memory · GitHub](https://github.com/EGO4D/episodic-memory/pull/42/files#diff-02c72650a457400be5a39796a6909311178f486289e2254a9c1b37c274ad8e05). EgoTracks/tools/preprocess/extract\_ego4d\_clip\_annotated\_frames.py only extracts annotated frames.

---

<div class="post-metadata">

**Author:** ![aram](https://avatars.discourse-cdn.com/v4/letter/a/a183cd/32.png) [@aram](https://discuss.ego4d-data.org/u/aram)\
**Post date:** [March 19, 2023, 9:22am UTC](https://discuss.ego4d-data.org/t/egotracks-dataset-download-failure/218/20 "2023-03-19T09:22:07Z")

</div>

Start processing db211359-c259-4515-9d6c-be521711b6d0!  
Start processing 87b52dc5-3ac3-47e7-9648-1b719049732f!  
Start processing b7fc5f98-e5d5-405d-8561-68cbefa75106!  
Start processing 59daca91-5433-48a4-92fc-422b406b551f!

I have problems preprocessing the videos above. (Process never ends and gives error)

```auto
File "av/enum.pyx", line 60, in av.enum.EnumType. __getitem__
KeyError: 'ERRORTYPE_2'

```

Could you please help me check what is going wrong with these?

[Next page](https://discuss.ego4d-data.org/t/egotracks-dataset-download-failure/218.md?page=2)
