File size: 2,120 Bytes
95456ed
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
import pandas as pd
import numpy as np
from tqdm import tqdm
import os
import random
import torch
from torch.utils.data import Dataset, DataLoader
from typing import Dict, Union, Any, Optional, List


class TextDataset(Dataset):

    def __init__(

        self,

        dataset,

        data_args,

        split,  # 'train', 'test', 'val'

    ):
        super().__init__()
        self.dataset = dataset
        self.length = len(self.dataset[split])
        self.data_args = data_args
        self.split = split

    def __len__(self):
        return self.length

    def __getitem__(self, idx):
        sample = {
            "mask": np.array(self.dataset[self.split][idx]["mask"]),
            "sn_sp_repr": np.array(self.dataset[self.split][idx]["sn_sp_repr"]),
            "sn_input_ids": np.array(self.dataset[self.split][idx]["sn_input_ids"]),
            "indices_pos_enc": np.array(self.dataset[self.split][idx]["indices_pos_enc"]),
            "sn_repr_len": np.array(self.dataset[self.split][idx]["sn_repr_len"]),
            "words_for_mapping": self.dataset[self.split][idx]["words_for_mapping"],
        }
        return sample


def text_dataset_loader(

    data,

    data_args,

    split: str,

    deterministic: bool = False,

):
    dataset = TextDataset(
        dataset=data,
        data_args=data_args,
        split=split,
    )
    data_loader = DataLoader(
        dataset,
        batch_size=data_args.batch_size,
        shuffle=not deterministic,
        num_workers=0,
    )
    return iter(data_loader)


class FixdurDataset(Dataset):
    def __init__(

        self,

        data: Dict[str, Union[torch.Tensor, Any]],

    ):
        super().__init__()
        self.data = data

    def __len__(self):
        return len(self.data["sp_embeddings"])

    def __getitem__(self, idx):
        sample = {
            "sp_embeddings": self.data["sp_embeddings"][idx],
            "attention_mask": self.data["attention_mask"][idx],
            "unique_idx": self.data["unique_idx"][idx],
        }
        return sample