KASKOWPusat Info

Cara mengambil teks dari file PDF tanpa mengetik ulangHow to extract text from a PDF without retyping

Ingin menyalin isi PDF ke Word, spreadsheet, atau catatan tapi hasil salin-tempel berantakan? Begini cara mengambil teksnya dan batasan yang perlu diketahui.

Want to copy a PDF's content into Word, a spreadsheet or notes but copy-paste comes out messy? Here is how to pull the text and the limits to know.

Diterbitkan 2 menit baca
Di halaman ini
  1. Dua jenis PDF, dua hasil berbeda
  2. Langkah mengambil teks
  3. Tips merapikan hasil
  4. Bagaimana dengan PDF hasil scan?
  1. Two kinds of PDF, two different results
  2. Steps to extract text
  3. Tips for cleaning the result
  4. What about scanned PDFs?

Mengetik ulang isi PDF itu lambat dan rawan salah. Padahal sering kali teks di dalamnya bisa diambil langsung.

Dua jenis PDF, dua hasil berbeda

Jenis PDFCiriBisa diambil teksnya?
PDF teksDibuat dari Word, laporan sistem, atau ekspor aplikasi. Teks bisa diblok dengan kursorYa
PDF gambar (hasil scan)Setiap halaman sebenarnya foto. Teks tidak bisa diblokTidak langsung, perlu pengenalan teks (OCR)

Cara cepat mengeceknya: coba blok satu kalimat dengan kursor. Jika bisa disorot dan disalin, itu PDF teks.

Langkah mengambil teks

  1. Buka alat PDF dan pilih fitur Ekstrak teks.
  2. Unggah file PDF.
  3. Jalankan, lalu salin atau unduh hasil teksnya.
  4. Rapikan seperlunya di editor: paragraf yang terpotong, kolom yang tercampur, dan nomor halaman.
Catatan

PDF dengan banyak kolom, tabel, atau tata letak kompleks bisa menghasilkan teks yang urutannya tidak sempurna. Periksa hasilnya, terutama angka.

Tips merapikan hasil

  • Hapus baris kosong ganda dengan fitur cari-ganti.
  • Untuk tabel, tempel hasilnya ke spreadsheet lalu pisahkan kolom secara manual.
  • Cek ulang angka penting seperti nominal, tanggal, dan nomor rekening. Jangan percaya begitu saja.

Bagaimana dengan PDF hasil scan?

Teks pada PDF scan perlu dikenali dengan OCR, dan hasilnya sangat bergantung pada kualitas pindaian. Ekstrak teks biasa tidak akan menemukan apa pun pada file semacam ini. Jika Anda membuat scan sendiri, pindai dengan resolusi cukup dan posisi lurus agar hasilnya lebih mudah diolah.

Fitur ekstrak teks tersedia gratis dan tanpa login di KASKOW PDF Converter. Fitur lainnya, seperti gabung dan kompres, ada di tempat yang sama.

Retyping a PDF's contents is slow and error-prone. Yet the text inside can often be pulled out directly.

Two kinds of PDF, two different results

PDF typeTraitsText extractable?
Text PDFMade from Word, system reports or app exports. Text can be highlighted with the cursorYes
Image PDF (scan)Each page is really a photo. Text cannot be highlightedNot directly; needs text recognition (OCR)

A quick check: try highlighting a sentence with the cursor. If it can be selected and copied, it is a text PDF.

Steps to extract text

  1. Open a PDF tool and choose Extract text.
  2. Upload the PDF.
  3. Run it, then copy or download the text result.
  4. Tidy up as needed in an editor: broken paragraphs, mixed columns and page numbers.
Note

PDFs with many columns, tables or complex layouts can yield text in an imperfect order. Check the result, especially numbers.

Tips for cleaning the result

  • Remove doubled blank lines with find-and-replace.
  • For tables, paste the result into a spreadsheet and split columns manually.
  • Re-check important figures such as amounts, dates and account numbers. Do not trust them blindly.

What about scanned PDFs?

Text in a scanned PDF needs OCR, and the outcome depends heavily on scan quality. Plain text extraction will find nothing in such files. If you make the scan yourself, scan at adequate resolution and straight so it is easier to process.

Text extraction is free and needs no login in KASKOW PDF Converter. Other features, like merge and compress, are in the same place.